Underwater optical image real-time enhancement method, system and device and storage medium

By constructing a homogeneous multi-stage image enhancement network and combining lightweight feature extraction and multi-scale fusion, the problem of high computational complexity of underwater image enhancement methods is solved, real-time underwater image enhancement is achieved on devices with limited computing resources, and the image processing speed and quality are improved.

CN120672619APending Publication Date: 2025-09-19CHINA YANGTZE POWER
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510598123.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods have high computational complexity, large resource consumption, slow processing speed and poor image quality, making it difficult to achieve real-time processing on devices with limited computing power.

Method used

A homogeneous multi-stage image enhancement network is constructed, point convolution and star convolution modules are used for feature extraction, multi-scale fusion and pixel attention mechanism are combined, lightweight feature extraction modules are used to reduce computational complexity, and the image restoration effect is optimized through the loss function.

Benefits of technology

It achieves real-time enhancement of underwater images on devices with limited computing resources, improves image processing speed and quality, and enhances the reliability of hydropower station inspection and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672619A_ABST
    Figure CN120672619A_ABST
Patent Text Reader

Abstract

The invention provides an underwater optical image real-time enhancement method, system and device and a storage medium, and relates to the technical field of image processing and recognition. An isomorphic multi-stage image enhancement network is constructed, the image enhancement network comprises four stages, four branches with different feature scales are obtained through down-sampling operation in each stage, and after each branch is connected with a corresponding star convolution module for feature extraction, the features are input into a multi-scale fusion module; after fusion, after extraction of a pixel attention mechanism module, a convolutional layer is spliced to output an enhanced image, progressive enhancement is realized through a cross-stage feature feedback mechanism, the enhanced image and features output by the later stage and the previous stage are used as input, underwater image recovery is performed, and multi-scale feature extraction is adopted to obtain an underwater image. The calculation complexity is remarkably reduced, the network degradation phenomenon caused by multi-stage training is avoided, and the reliability of efficient detection and maintenance of the hydropower station is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and recognition, and in particular to a method, system, device and storage medium for real-time enhancement of underwater optical images. Background Art

[0002] Surface power platforms play a vital role in the inspection and maintenance of hydropower stations. Their onboard optical imaging systems can capture underwater images of dam structures, providing critical data support for subsequent maintenance and repair. However, underwater images often exhibit blurring due to light scattering by suspended particles in the water. Furthermore, because red light attenuates more rapidly than green and blue light underwater, underwater images often suffer from color casts and other issues.

[0003] Traditional underwater image enhancement methods are primarily based on physical models or image processing techniques, such as histogram equalization, white balance adjustment, and dehazing algorithms based on dark channel priors. While these methods can improve image quality to a certain extent, they often rely on artificially designed features and assumptions, making them difficult to adapt to the complex and changing underwater environment. Furthermore, these methods suffer from high computational complexity, making them difficult to meet real-time processing requirements. This is particularly true when running on resource-constrained devices, such as mobile surface platforms, drones, or portable devices, where performance is particularly limited.

[0004] In recent years, the rapid development of deep learning technology has provided new solutions for underwater image enhancement. Methods based on convolutional neural networks (CNNs) and generative adversarial networks (GANs) have achieved remarkable results in image enhancement tasks, effectively restoring the color, contrast, and detail information of underwater images. However, these deep learning methods typically require a large amount of computing resources and storage space. Their complex network structure and large number of parameters make them difficult to run efficiently on devices with limited computing power and storage space. For example, surface power platforms or drones are often equipped with low-power computing units that cannot support real-time inference of large-scale neural networks. Therefore, how to reduce the computational complexity and resource consumption of the algorithm while ensuring the image enhancement effect has become a pressing issue. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, system, device and storage medium for real-time enhancement of underwater optical images, so as to solve the technical problems in the prior art such as insufficient image processing accuracy, high algorithm computational complexity, slow processing speed and poor image quality.

[0006] To solve the above technical problems, the present invention adopts a technical solution: a method for real-time enhancement of underwater optical images, comprising the following steps: S1: The optical imaging system of the surface power platform takes real-time video stream data of the underwater environment; S2: Collect the video stream data of the underwater environment and extract features through a point convolution; S3: Construct a homogeneous basic stage network, input the feature data extracted by point convolution, and output an underwater enhanced image; the basic stage network includes four branches, each of which obtains four different feature maps through downsampling operations. Each feature map is connected to the corresponding star convolution module. After the downsampled feature maps of different branches are subjected to feature extraction through star convolution, feature maps of different sizes at the current scale are output. While being provided to the next stage, they are also input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module and spliced ​​into a convolution layer to output the enhanced image; S4: After obtaining the output enhanced image and feature map of S3 according to the preset four stages, it is used as the input of the next stage and step S3 is repeated three times to finally obtain the output map of the four stages; S5: Use loss function and image enhancement network for training; S6: Based on the computing power of the deployment platform and the preset indicator requirements, perform structured pruning on the optimal image enhancement network and save the optimal image enhancement network; S7: Input the video stream data after the point convolution to be tested, use the optimal image enhancement network for inference, and output the underwater enhanced image.

[0007] In the preferred solution, the video stream data is obtained, and after data preprocessing, the input image is scaled to Resolution size.

[0008] In the preferred solution, in step S3, the four downsampling modules of different scales operate as follows: ; ; ; ; in, is the input feature map, For the stage, Indicates progress The average pooling operation of downsampling is For the previous stage branches, For the stage.

[0009] In the preferred solution, in step S3, the star convolution module specifically operates as follows: First input feature map , after convolution and instance normalization operations: ; in, For instance normalization, It is a convolution operation with a convolution kernel of 3; Secondly, the obtained feature map is subjected to parallel point convolution and instance normalization operations: ; ; in, It is a point convolution operation with a convolution kernel of 1; Again, the two parallel feature maps are fused element-by-element by multiplication: ; in, It is a point convolution operation with a convolution kernel of 1; Finally, the feature fusion is performed through point convolution and instance normalization to output the final feature : .

[0010] In the preferred solution, in step S3, the fusion module specifically operates as follows: First, input multi-scale features ,in is a feature map with the same resolution as the original image, The feature maps are 8, 16 and 32 downsampling times respectively; firstly, each feature is nonlinearly activated and the low-resolution features are upsampled to make their resolution equal to Consistent: ; ; ; ; in, Represents upsampled feature maps of different resolutions, represents the linear rectification function, Represents the upsampling operation, which adjusts the feature resolution to size; Then add the upsampled features to the highest resolution features for feature fusion: ; The fused features are convolved to generate the final output features: ; in, It is a convolution operation with a convolution kernel of 3.

[0011] In a preferred embodiment, in S5, the loss function includes: Reconstruction loss function , the formula is: ; in, express loss function, represents the structural consistency loss function, Indicates a reference image, Indicates the Stage restored image; Perceptual loss function , the formula is: ; in, Represents pre-trained Feature maps extracted by the encoder; Performance improvement loss function, the formula is: ; Dynamic enhancement loss function, the formula is: ; ; ; ; ; in, is the dynamic enhancement loss, For the iterations, is the total number of training iterations, is the dynamic enhancement factor.

[0012] In the preferred solution, in step S6, the structured pruning is specifically as follows: according to the computing power and speed requirements of different platforms, the deployment In this stage, the subsequent network is deleted and the obtained network is inferred.

[0013] A real-time underwater optical image enhancement system, applicable to a real-time underwater optical image enhancement method, comprising: The data acquisition and processing module is used to use the optical imaging system of the surface power platform to collect video stream data of the underwater environment in real time, and extract features from the collected video stream data of the underwater environment through a point convolution; A network module is constructed to construct an isomorphic basic stage network, input feature data extracted by point convolution, and output an underwater enhanced image. The basic stage network includes four branches, each of which obtains four different feature maps through downsampling operations. Each feature map is connected to a corresponding star convolution module. After the downsampled feature maps of different branches are subjected to feature extraction through star convolution, feature maps of different sizes at the current scale are output. These feature maps are provided to the next stage and then input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module and then spliced ​​into a convolution layer to output the enhanced image. The multi-stage module is used to obtain the output map and feature map of S3 according to the preset four stages, and use them as the input of the next stage. Step S3 is repeated three times to finally obtain the output map of four stages; Network training module, used to train image enhancement network using loss function; The pruning module is used to perform structured pruning on the optimal image enhancement network based on the computing power of the deployment platform and preset indicator requirements, and save the optimal image enhancement network; The application module is used to input video stream data after convolution at a point to be tested, use the optimal image enhancement network for inference, and output underwater enhanced images.

[0014] An electronic device comprising a memory and a processor; The memory is used to store computer programs; The processor is used to implement the method for real-time enhancement of underwater optical images when executing the computer program.

[0015] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for real-time enhancement of underwater optical images is implemented.

[0016] The present invention provides a real-time enhancement method for underwater optical images. The method constructs an isomorphic multi-stage image enhancement network, which includes four stages. The latter stage uses the enhanced image and features output by the previous stage as input, connects to the corresponding star convolution module for feature extraction, and then inputs the feature into a multi-scale fusion module. After fusion, the image is extracted by a pixel attention mechanism module, and then a convolution layer is spliced ​​to output the enhanced image, thereby realizing a cross-stage feature feedback mechanism to achieve progressive enhancement. That is, the latter stage uses the enhanced image and features output by the previous stage as input to restore the underwater image. Secondly, a multi-branch parallel structure is adopted to realize multi-scale feature extraction, and a lightweight feature extraction module and a multi-scale feature fusion module are combined to perform feature fusion, which significantly reduces the computational complexity while ensuring the feature expression capability. Finally, the skip connection and the designed loss function are utilized to effectively alleviate the degradation phenomenon of the network generated by multi-stage training, solve the technical problems of slow image processing speed and poor quality of surface power platform images, and improve the reliability of efficient detection and maintenance of hydropower stations.

[0017] Furthermore, the loss functions adopted in the present invention include reconstruction loss function, perception loss function, performance improvement loss function and dynamic enhancement loss function, which optimize the point-by-point alignment at the pixel level. By minimizing the pixel difference between the restored image and the reference image, the fidelity of the image in the spatial domain is ensured, the semantic consistency of the image is improved, and the restoration performance of the model is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will be further described below with reference to the accompanying drawings and examples: Figure 1 It is a model structure diagram of the present invention; Figure 2 It is a basic stage structure diagram of the present invention; Figure 3 It is the basic block and feature fusion module of the present invention; Figure 4 The Nvidia Orin Nx 16G embedded development board of the present invention; Figure 5 This is a comparison diagram of the original image and the final enhanced image output of the present invention. DETAILED DESCRIPTION

[0019] Example 1 like Figure 1-5 As shown, a real-time underwater optical image enhancement method includes the following steps: S1: The optical imaging system of the surface power platform takes real-time video streaming data of the underwater environment.

[0020] S2: Collect the video stream data of the underwater environment and extract features through a point convolution: ; in, represents the convolution operation, is the original underwater image with a resolution of , The output feature map resolution is .

[0021] S3: Construct an isomorphic basic stage network, input the feature data extracted by point convolution, and output the underwater enhanced image; the basic stage network includes four branches, each branch obtains four different feature maps through downsampling operation, each feature map is connected to the corresponding star convolution module, and the downsampled feature maps of different branches are subjected to feature extraction through star convolution, and the feature maps of different sizes of the current scale are output. While being provided to the next stage for use, they are then input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module and spliced ​​into a convolution layer to output the enhanced image.

[0022] S4: After obtaining the output enhanced image and feature map of S3 according to the preset number of 4 stages, it is used as the input of the next stage and step S3 is repeated 3 times to finally obtain the output map of the 4 stages, specifically: First, the output map of the previous stage is concatenated with the output feature map of the previous stage at the same image resolution. Point convolution is then performed to obtain an initial feature map. This initial feature map is downsampled and added to the feature map of the current stage at the same downsampled scale to obtain feature maps of four different branch scales. Each feature map is connected to the corresponding star convolution module. Star convolution is used to extract features from the downsampled feature maps of different branches, and feature maps of different sizes at the current scale are output. These feature maps are provided to the next stage and then input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module, and then a convolution layer is spliced ​​to output the enhanced image. This step is repeated three times to obtain the output maps of the four stages.

[0023] S5: Using loss function, image enhancement network is trained.

[0024] S6: Based on the computing power of the deployment platform and the preset indicator requirements, perform structured pruning on the optimal image enhancement network and save the optimal image enhancement network.

[0025] S7: Input the video stream data after the point convolution to be tested, use the optimal image enhancement network for inference, and output the underwater enhanced image.

[0026] This embodiment is based on a multi-stage optimization architecture of a homogeneous network to construct a homogeneous multi-stage image enhancement network. The image enhancement network includes four stages, and the latter stage takes the enhanced image and features output by the previous stage as input, connects to the corresponding star convolution module for feature extraction, and then inputs it into the multi-scale fusion module. After fusion, it is extracted by the pixel attention mechanism module, and then a convolution layer is spliced ​​to output the enhanced image, realizing a cross-stage feature feedback mechanism to achieve progressive enhancement, that is, the latter stage takes the enhanced image and features output by the previous stage as input to restore the underwater image; secondly, a multi-branch parallel structure is adopted to realize multi-scale feature extraction, and a lightweight feature extraction module and a multi-scale feature fusion module are combined for feature fusion, which significantly reduces the computational complexity while ensuring the feature expression ability; finally, the use of jump connections and designed loss functions effectively alleviates the network degradation phenomenon caused by multi-stage training, solves the technical problems of slow speed and poor quality of image processing of surface power platforms, and improves the reliability of efficient detection and maintenance of hydropower stations.

[0027] like Figure 1 FIG. 4 is a schematic diagram of a network implementation of this embodiment.

[0028] This embodiment utilizes the optical imaging system of the surface power platform to acquire video stream data of the underwater environment in real time.

[0029] In this embodiment, the model adopts a structural consistency design. The network structure of each stage is the same and can independently output enhancement results. Through structured pruning technology, the number of deployment stages can be flexibly selected according to the computing power of the deployment platform. The algorithm of the present invention adopts the first-stage deployment and can achieve real-time processing of 720p FPS21+ on the embedded platform NVIDIA Jetson Orin NX 16GB, fully meeting the deployment requirements in edge computing scenarios.

[0030] In the preferred solution, the video stream data is obtained, and after data preprocessing, the input image is scaled to Resolution size.

[0031] Through structured pruning technology, the number of deployment stages can be adaptively selected based on the computing power and speed requirements of different deployment platforms.

[0032] like Figure 2 As shown in FIG, the basic stage structure diagram of this embodiment is shown, in which the basic block is the star convolution module, where: Figure 2 (a) is the network structure diagram of the first stage. The feature data extracted by point convolution is input from the left and enters the isomorphic basic stage network with four branches.

[0033] Figure 2(b) shows the network structure for stages 2-4. Compared to stage 1, connections to the previous stage are added. The output map of the previous stage is concatenated with the output feature map of the same resolution. After pointwise convolution and downsampling, it is added to the downsampled feature map of the current stage to generate four new feature maps at different branch scales, enabling cross-stage feature feedback. The subsequent process is similar to that of stage 1. The newly generated feature map is processed through a star convolution module, a multi-scale fusion module, a pixel attention mechanism module, and a convolutional layer to output the enhanced image of the corresponding stage. Each stage optimizes the results of the previous stage, gradually improving image quality, ultimately completing the progressive enhancement of underwater optical images.

[0034] Step S3: Input the homogeneous multi-stage underwater image enhancement network for processing, including the following process: S301: The downsampling modules of four different scales are operated as follows: Enter Feature map When , downsampling is performed to different sizes by 8, 16, or 32 times to obtain feature maps of different resolutions: ; ; ; ; in, is the input feature map, For the stage, Indicates progress The average pooling operation of downsampling is For the previous stage branches, For the stage.

[0035] S302: The specific operations of the star convolution module are as follows: The downsampled feature maps of different branches are subjected to star convolution for feature extraction to obtain accurate underwater features, and feature maps of different sizes of the current scale are output for use in the next stage.

[0036] The basic block module is operated as follows: first input Feature map , after convolution and instance normalization operations: ; in, For instance normalization, It is a convolution operation with a convolution kernel of 3.

[0037] Secondly, the obtained feature map is subjected to parallel point convolution and instance normalization operations: ; ; in, It is a point convolution operation with a convolution kernel of 1.

[0038] Again, the two parallel feature maps are fused element-by-element by multiplication: ; in, It is a point convolution operation with a convolution kernel of 1.

[0039] Finally, the feature fusion is performed through point convolution and instance normalization to output the final feature

[0040] .

[0041] S303: Input the underwater features extracted from different branches into the fusion module for fusion: First, the underwater features of different branches are input into the multi-scale feature fusion module to obtain multi-scale features. ,in yes The feature map has the same resolution as the original image. The feature maps are 8, 16 and 32 downsampling times respectively; firstly, each feature is nonlinearly activated and the low-resolution features are upsampled to make their resolution equal to Consistent: ; ; ; ; in, Represents upsampled feature maps of different resolutions, represents the linear rectification function, Represents the upsampling operation, which adjusts the feature resolution to size.

[0042] Then, the upsampled features are added to the highest resolution features for feature fusion: .

[0043] The fused features are convolved to generate the final output features: ; in, It is a convolution operation with a convolution kernel of 3.

[0044] S304: The fusion module is passed through the pixel attention mechanism module to extract important information.

[0045] S305: Output the enhanced image through the convolution layer.

[0046] Through the above steps, this embodiment adopts a multi-branch parallel structure to perform multi-scale feature extraction, combined with a lightweight star convolution module and a multi-scale feature fusion module, while ensuring the feature expression capability, it significantly reduces the computational complexity.

[0047] This embodiment trains the network using the designed loss function, which includes reconstruction loss function, perception loss function, performance improvement loss function and dynamic enhancement loss function. Among them, the reconstruction loss function mainly emphasizes the point-by-point alignment at the pixel level, and ensures the fidelity of the image in the spatial domain by minimizing the pixel difference between the restored image and the reference image; the perception loss function is based on the high-level semantic features extracted by the pre-trained VGG19, and improves the semantic consistency of the image by minimizing the distance between the restored image and the reference image in the feature space. In addition, if the multi-stage restoration model proposed in the present invention is not restricted, the network may degenerate, that is, the restoration effect of the latter stage is worse than that of the previous stage. To this end, the introduction of the performance improvement loss function and the dynamic enhancement loss function effectively alleviates this problem, ensuring that the restoration result of the latter stage is always better than that of the previous stage, thereby gradually approaching the reference image.

[0048] In the preferred solution, in step S5, the loss function used includes: Reconstruction loss function , the formula is: ; in, express loss function, represents the structural consistency loss function, Indicates a reference image, Indicates the Stage restored image.

[0049] Perceptual loss function , the formula is: ; in, Represents pre-trained Feature maps extracted by the encoder.

[0050] Performance improvement loss function, the formula is: ; Dynamic enhancement loss function, the formula is: ; ; ; ; ; in, is the dynamic enhancement loss, For the iterations, is the total number of training iterations, is the dynamic enhancement factor.

[0051] This implementation utilizes skip connections and a designed composite loss function consisting of reconstruction, perception, performance improvement, and dynamic enhancement to address network degradation caused by multi-stage training. The use of the performance improvement and dynamic enhancement losses ensures that the restoration results of each subsequent stage are consistently better than those of the previous one, gradually approaching the reference image and improving the model's overall restoration performance.

[0052] In the preferred solution, in step S6, the structured pruning is specifically: Determine the computing power and speed requirements of different platforms before deployment. In this stage, the subsequent network is deleted and the obtained network is inferred.

[0053] In this embodiment, the entire network is deployed on the RTX3090 graphics card, and one stage is deployed on the embedded device. Real-time inference can be performed on different platforms according to the recovery quality requirements.

[0054] In step S7, the pruned model is used for inference to output an underwater enhanced image.

[0055] like Figure 5 As shown in FIG. 1 , a comparison diagram of the original image and the underwater enhanced image output by this embodiment is shown. It can be seen that the image restoration effect is continuously optimized through the above steps, the processing accuracy of the underwater optical image is improved, the clarity of the image details is improved, and the color accuracy is improved.

[0056] This embodiment uses the official nvidia tool to deploy the adopted algorithm to the Orin Nx development board, such as Figure 4As shown in the figure, the specific process is as follows: First, export the trained model to ONNX format; second, optimize and convert the model through TensorRT tools to generate an efficient TensorRT engine and generate an .engine file; then, load the .engine file in the inference phase, prepare input data, perform inference, and obtain output results.

[0057] The network structure used in this embodiment reduces computational complexity and enables real-time processing of underwater optical images. Using the first phase of deployment on the NVIDIA Jetson Orin NX 16GB embedded platform, it can achieve real-time processing at 21+ FPS at 720p resolution, greatly improving the image processing speed of surface power platforms.

[0058] Test by deploying TensorRT to improve the inference efficiency of deep learning models on NVIDIA GPUs to maximize the performance of embedded platforms.

[0059] Example 2 In conjunction with Example 1, a real-time underwater optical image enhancement system based on a homogeneous network is provided, which is applicable to a real-time underwater optical image enhancement method, including: The data acquisition and processing module is used to use the optical imaging system of the surface power platform to collect video stream data of the underwater environment in real time, and extract features from the collected video stream data of the underwater environment through a point convolution.

[0060] A network module is constructed to construct an isomorphic basic stage network, input feature data extracted by point convolution, and output an underwater enhanced image; the basic stage network includes four branches, each branch obtains four different feature maps through downsampling operation, each feature map is connected to the corresponding star convolution module, and the downsampled feature maps of different branches are subjected to feature extraction through star convolution, and feature maps of different sizes of the current scale are output. While being provided to the next stage for use, they are then input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module, and then a convolution layer is spliced ​​to output the enhanced image.

[0061] The multi-stage module is used to first concatenate the output map of the previous stage with the output feature map of the previous stage with the same image resolution according to the preset number of four stages. Then, the initial feature map is obtained through point convolution. The obtained initial feature map is downsampled and added to the feature map of the same downsampled scale in the current stage to obtain feature maps of four different branch scales. Each feature map is connected to the corresponding star convolution module. After the downsampled feature maps of different branches are extracted through star convolution, the feature maps of different sizes of the current scale are output. They are provided to the next stage and then input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module and a convolution layer is spliced ​​to output the enhanced image. This step is repeated three times to finally obtain the output map of the four stages.

[0062] The network training module is used to train the image enhancement network using the loss function.

[0063] The pruning module is used to perform structured pruning on the optimal image enhancement network according to the computing power of the deployment platform and preset indicator requirements, and save the optimal image enhancement network.

[0064] The application module is used to input video stream data after convolution at a point to be tested, use the optimal image enhancement network for inference, and output underwater enhanced images.

[0065] An electronic device comprising a memory and a processor; Memory for storing computer programs.

[0066] The processor is configured to implement a real-time underwater optical image enhancement method as described in Example 1 when executing the computer program.

[0067] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, a real-time underwater optical image enhancement method as described in Example 1 is implemented.

[0068] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A real-time underwater optical image enhancement method, characterized in that: The following steps are involved: S1: The optical imaging system of the surface power platform takes real-time video stream data of the underwater environment; S2: Collect the video stream data of the underwater environment and extract features through a point convolution; S3: Construct a homogeneous basic stage network, input the feature data extracted by point convolution, and output an underwater enhanced image; the basic stage network includes four branches, each of which obtains four different feature maps through downsampling operations. Each feature map is connected to the corresponding star convolution module. After the downsampled feature maps of different branches are subjected to feature extraction through star convolution, feature maps of different sizes at the current scale are output. While being provided to the next stage, they are also input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module and spliced ​​into a convolution layer to output the enhanced image; S4: After obtaining the output enhanced image and feature map of S3 according to the preset four stages, it is used as the input of the next stage and step S3 is repeated three times to finally obtain the output map of the four stages; S5: Use loss function and image enhancement network for training; S6: Based on the computing power of the deployment platform and the preset indicator requirements, perform structured pruning on the optimal image enhancement network and save the optimal image enhancement network; S7: Input the video stream data after the point convolution to be tested, use the optimal image enhancement network for inference, and output the underwater enhanced image.

2. The method for real-time enhancement of underwater optical images according to claim 1, characterized in that: Get the video stream data, pre-process the data, and scale the input image to Resolution size.

3. The method for real-time enhancement of underwater optical images according to claim 1, characterized in that: In step S3, the four downsampling modules of different scales operate as follows: ; ; ; ; in, is the input feature map, For the stage, Indicates progress The average pooling operation of downsampling is For the previous stage branches, For the stage.

4. The method for real-time enhancement of underwater optical images according to claim 1, characterized in that: In step S3, the star convolution module specifically operates as follows: First input feature map , after convolution and instance normalization operations: ; in, For instance normalization, It is a convolution operation with a convolution kernel of 3; Secondly, the obtained feature map is subjected to parallel point convolution and instance normalization operations: ; ; in, It is a point convolution operation with a convolution kernel of 1; Again, the two parallel feature maps are fused element-by-element by multiplication: ; in, It is a point convolution operation with a convolution kernel of 1; Finally, the feature fusion is performed through point convolution and instance normalization to output the final feature : 。 5. The method for real-time enhancement of underwater optical images according to claim 1, characterized in that: In step S3, the fusion module performs the following operations: First, input multi-scale features ,in is a feature map with the same resolution as the original image, The feature maps are 8, 16 and 32 downsampling times respectively; firstly, each feature is nonlinearly activated and the low-resolution features are upsampled to make their resolution equal to Consistent: ; ; ; ; in, Represents upsampled feature maps of different resolutions, represents the linear rectification function, Represents the upsampling operation, which adjusts the feature resolution to size; Then add the upsampled features to the highest resolution features for feature fusion: ; The fused features are convolved to generate the final output features: ; in, It is a convolution operation with a convolution kernel of 3.

6. The method for real-time enhancement of underwater optical images according to claim 1, characterized in that: In S5, the loss function includes: Reconstruction loss function , the formula is: ; in, express loss function, represents the structural consistency loss function, Indicates a reference image, Indicates the Stage restored image; Perceptual loss function , the formula is: ; in, Represents pre-trained Feature maps extracted by the encoder; Performance improvement loss function, the formula is: ; Dynamic enhancement loss function, the formula is: ; ; ; ; ; in, is the dynamic enhancement loss, For the iterations, is the total number of training iterations, is the dynamic enhancement factor.

7. The method for real-time enhancement of underwater optical images according to claim 1, characterized in that: In step S6, the structured pruning is specifically as follows: according to the computing power and speed requirements of different platforms, the deployment In this stage, the subsequent network is deleted and the obtained network is inferred.

8. A real-time underwater optical image enhancement system, characterized in that: A real-time underwater optical image enhancement method is provided, comprising: The data acquisition and processing module is used to use the optical imaging system of the surface power platform to collect video stream data of the underwater environment in real time, and extract features from the collected video stream data of the underwater environment through a point convolution; A network module is constructed to construct an isomorphic basic stage network, input feature data extracted by point convolution, and output an underwater enhanced image. The basic stage network includes four branches, each of which obtains four different feature maps through downsampling operations. Each feature map is connected to a corresponding star convolution module. After the downsampled feature maps of different branches are subjected to feature extraction through star convolution, feature maps of different sizes at the current scale are output. These feature maps are provided to the next stage and then input into the multi-scale fusion module. After fusion, they are extracted by the pixel attention mechanism module and then spliced ​​into a convolution layer to output the enhanced image. The multi-stage module is used to obtain the output map and feature map of S3 according to the preset four stages, and use them as the input of the next stage. Step S3 is repeated three times to finally obtain the output map of four stages; Network training module, used to train image enhancement network using loss function; The pruning module is used to perform structured pruning on the optimal image enhancement network based on the computing power of the deployment platform and preset indicator requirements, and save the optimal image enhancement network; The application module is used to input video stream data after convolution at a point to be tested, use the optimal image enhancement network for inference, and output underwater enhanced images.

9. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the method for real-time enhancement of underwater optical images according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the method for real-time enhancement of underwater optical images according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Construction method, system and equipment of integrated underwater image enhancement system and medium

    CN121563811A