General integrated water surface image enhancement and target detection system and method
By designing a universal integrated water surface image enhancement and object detection system, using generator networks and multi-scale enhancement discriminator networks, the problems of existing water surface object detection algorithms in dealing with diversified water surface targets are solved, and efficient image enhancement and object detection performance are achieved.
Patent Information
- Application Number
- CN202510117199.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
AI Technical Summary
When existing surface object detection algorithms deal with diversified surface targets, it is difficult to deal with problems such as the motion blur of the target. The traditional methods are complex, and deep learning methods have insufficient quality and efficiency.
A universal integrated water surface image enhancement and object detection system is designed. Through the generator network and multi-scale enhancement discriminator network, the image enhancement oriented by object detection is realized and the performance of water surface object detection is improved.
While ensuring image enhancement quality, the system improves image enhancement efficiency, improves the performance and real-time performance of surface target detection, and can better handle surface ship targets of different aspect ratios.
Smart Images

Figure CN120070216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image enhancement and object detection, and particularly to a general integrated water surface image enhancement and object detection system and method. Background Art
[0002] The main task of object detection is to classify and locate objects in an image, which is the basis for tasks such as object tracking and image segmentation. Object detection based on visible light images of the water surface is one of the essential key technologies for ships to achieve intelligent navigation.
[0003] Existing object detection algorithms are mainly divided into two categories: traditional methods and deep learning-based methods. Traditional water surface object detection algorithms must design many sliding windows of different sizes and aspect ratios, which greatly increases the complexity of the algorithm. Moreover, traditional image feature extraction algorithms mainly rely on handcrafted features and are stretched when dealing with diverse water surface object detections, and are difficult to handle problems such as motion blur of objects. Object detection algorithms based on convolutional neural networks can be roughly divided into two-stage methods and one-stage methods. One-stage methods do not need to generate candidate regions during the detection process, but directly predict the position and category of objects based on pre-set fixed-size anchor boxes. Two-stage methods have relatively high detection quality but are slow, while one-stage methods have more advantages in inference efficiency but slightly lower quality.
[0004] Therefore, it is necessary to invent a general integrated water surface image enhancement and object detection system and device that can ensure the quality of image enhancement while taking into account the efficiency of image enhancement, and highly correlate image enhancement with object detection to achieve image enhancement guided by object detection, thereby improving the performance of water surface object detection. Summary of the Invention
[0005] The purpose of the present invention is to provide a general integrated water surface image enhancement and object detection system, which can ensure the quality of image enhancement while taking into account the efficiency of image enhancement, achieve image enhancement guided by object detection, and improve the performance of water surface object detection.
[0006] To achieve this purpose, a general integrated water surface image enhancement and object detection system designed by the present invention includes:
[0007] An image acquisition module is used to acquire clear water surface images and corresponding original water surface images to be enhanced at each water surface position;
[0008] An image restoration module is used to input the original water surface images to be enhanced at each water surface position into a generator network model to generate initial restored images at each water surface position;
[0009] The image discrimination module is used to input the initial restored images and the corresponding clear images at each water surface position into the multi-scale enhanced discriminator network model for discriminating the clarity of the initial restored images at each water surface position, and updating the model parameters in the generator network model according to the clarity discrimination results to obtain the water surface image enhancement generator network model;
[0010] The detection and recognition module inputs the initial restored images and the corresponding clear images at each water surface position into the multi-aspect ratio detector network model to obtain the target detection and recognition results, and the generator network model updates the model parameters in the water surface image enhancement generator network model according to the target detection and recognition results to obtain an integrated generator network model that is conducive to water surface image enhancement and target detection of the restored images.
[0011] Preferably, the generator network model and the multi-scale enhanced discriminator network model are constructed based on the generative adversarial network. The generator network model is responsible for enhancing the original water surface image; the multi-aspect ratio detector network model is responsible for detecting the targets in the initial restored images; the multi-scale enhanced discriminator network model discriminates the quality of the restored images based on the backbone network of the multi-aspect ratio detector, so as to realize the integration of water surface image enhancement and target detection.
[0012] Preferably, the specific method for generating the initial restored images at each water surface position is as follows: input the image to be enhanced into the generator network model, obtain the preliminary features of the image to be enhanced through the feature extraction structure composed of multiple residual convolutional blocks and instance normalization layers, and the feature map is reconstructed to the original resolution of the image through the network structure composed of multiple deconvolution modules and feature extraction structures at the backend to obtain the initial restored images.
[0013] Preferably, the generator network model is reconstructed based on the generative adversarial network model to enhance the original water surface image. The front end of the network adopts a feature extraction structure composed of multiple residual convolutional blocks and instance normalization layers, which is responsible for extracting the preliminary features of the original water surface image to be enhanced; the middle part of the network adopts multiple receptive field enhancement modules to enhance the image feature extraction ability of the generator network model, and at the same time sets a global skip connection structure from the output of the front-end feature map to the middle receptive field enhancement module, enabling the generator network model to learn the residuals between the original water surface image to be enhanced and the clear image; the backend of the network uses a network structure composed of multiple deconvolution modules and feature extraction structures for image reconstruction, which is responsible for reconstructing the restored image to the original resolution.
[0014] Preferably, the receptive field enhancement module is divided into multiple branches. Each branch first uses convolution to perform dimensionality reduction on the input data. Based on the skip connection idea of the residual network, one of the branches after dimensionality reduction is directly connected to the front of the activation layer of the network. According to the network structure of Inception, the remaining branches respectively use convolution kernels of different sizes to perform convolution processing on the dimensionality-reduced features. The convolution kernels of different sizes can obtain receptive fields of different sizes, thereby enhancing the scale adaptability of the network to the input samples and capturing image information at different levels. At the ends of the branches, dilated convolution processing with different dilation rates is performed respectively to expand the receptive field of the network without reducing the image resolution. After the dilated convolution processing, the remaining branches are concatenated in the channel dimension, and then convolution processing is performed and fused with the data of the other branch directly connected to the front of the activation layer of the network. Finally, it is output through the ReLU activation layer.
[0015] Preferably, the multi-scale enhanced discriminator network model and the multi-aspect ratio detector network model share a backbone network. Three branches are led out from the last three residual units of the backbone network. Each branch uses three consecutive convolutional layers to compress the feature maps of the residual units and inputs them into the back-end networks of the multi-scale enhanced discriminator network and the multi-aspect ratio detector network respectively. Sharing the backbone network reduces the number of network parameters, enables the image quality discrimination and object detection to share the features in the backbone network, performs image quality discrimination at multiple scales, and then conducts adversarial training with the generator based on this.
[0016] Preferably, the front end of the multi-scale enhanced discriminator network model is the backbone network, which extracts features of the image from shallow to deep through consecutive convolutional blocks and residual blocks, and abandons the max-pooling operation. Three branches are led out from the last three residual units of the backbone network. Each branch uses three consecutive convolutional layers to compress the feature maps of the residual units into a single-channel two-dimensional matrix. Each element in the matrix represents the discrimination result of the image quality within its receptive field, and the average value of all elements in the matrix is used to reflect the overall quality of the image.
[0017] Preferably, the front end of the multi-aspect ratio detector network model is the shared backbone network for feature extraction. There are three branches in the back end. In each branch, the features are preliminarily processed using a set number of convolutional blocks. The deep image features after processing are compressed and upsampled through convolutional blocks and then concatenated with the shallow features. At the same time, the three branches are respectively used to process water surface ship targets with different aspect ratios. Among them, the scale of the convolutional kernel of the branch for processing deep features is set to 1×3; the scale of the convolutional kernel of the branch for processing shallow features is set to 3×5; the last branch uses a convolutional kernel size of 3×3 to process the detection of square ship targets.
[0018] Advantages of the present invention: The present invention proposes a general integrated water surface image enhancement and target detection system, which highly correlates water surface image enhancement and target detection. A receptive field module is introduced into the generator to strengthen the feature extraction ability of the generator for the original water surface image to be enhanced, and at the same time improve the image restoration efficiency of the network. A multi-scale enhanced discriminator is proposed, which can not only effectively reduce the number of network parameters, but also enable the quality discrimination of the image and the target detection to share the features in the backbone network. The quality of the restored image is discriminated based on the features related to target detection at multiple scales, helping the generator to restore an image that is more conducive to target detection. A detector head network is designed according to the shape characteristics of ship targets. Multiple branches are respectively used to process water surface ship targets with different aspect ratios, more specifically detecting ship targets, and at the same time can reduce the number of parameters of the detector, improving the accuracy and real-time performance of the target detection system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic structural diagram of the present invention;
[0020] Figure 2 It is a schematic flow diagram of the present invention;
[0021] Figure 3 It is a schematic flow diagram of a single image;
[0022] Figure 4 It is a schematic diagram of the generator network model;
[0023] Figure 5 It is a schematic diagram of the multi-scale enhanced discriminator network model;
[0024] Figure 6 It is a schematic diagram of the multi-aspect ratio detector network model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0026] The following further elaborates on the present invention in detail in conjunction with the accompanying drawings and specific embodiments:
[0027] Embodiment 1
[0028] A general integrated water surface image enhancement and target detection system, as Figure 1 shown, it includes:
[0029] The image acquisition module is used to acquire clear water surface images and corresponding original water surface images to be enhanced at each water surface position;
[0030] The image restoration module is used to input the original water surface images to be enhanced at each water surface position into the generator network model to generate initial restored images at each water surface position;
[0031] The image discrimination module is used to input the initial restored images and corresponding clear images at each water surface position into the multi-scale enhancement discriminator network model for clarity discrimination of the initial restored images at each water surface position, and update the model parameters in the generator network model according to the clarity discrimination results to obtain the water surface image enhancement generator network model;
[0032] The detection and recognition module inputs the initial restored images and corresponding clear images at each water surface position into the multi-aspect ratio detector network model to obtain the target detection and recognition results. The generator network model updates the model parameters in the water surface image enhancement generator network model according to the target detection and recognition results to obtain an integrated generator network model that is conducive to water surface image enhancement and target detection of the restored images.
[0033] In the above technical solution, the targets in the target detection and recognition results are set as different targets according to specific tasks. The main target detected in the experiment is a ship. The corresponding ship classes are recognized and output according to the ship targets detected in the acquired images, such as cargo ships, oil tankers, passenger ships, etc. Among them, the detection target may also be a buoy or a navigation mark, etc.
[0034] In the above technical solution, the Seaships dataset is adopted, which contains 7000 pictures, including six types of ship targets such as container ships, bulk carriers, ore carriers, general cargo ships, fishing boats, and passenger ships. 5000 training images and 2000 validation set images are randomly divided. The image resolution in the experiment is 960×544. Motion blur is added to the clear images based on the motion blur physical model to obtain the corresponding original water surface images to be enhanced. At this time, there are paired clear images and corresponding original water surface images to be enhanced.
[0035] In the above technical solution, the specific training steps of the network model of the general integrated water surface image enhancement and target detection system and device are as follows: Randomly extract 7,000 images from the Seaships dataset, randomly divide 5,000 images as the training set and 2,000 images as the test set. Add motion blur to the clear images based on the motion blur physical model to obtain the corresponding original water surface images to be enhanced. At this time, there are paired clear images and corresponding original water surface images to be enhanced; Send the clear images in the training set into the network to train the backbone network and detection head network of the multi-aspect ratio detector. At this time, the backbone network trained by the clear ship images will pay more attention to the features related to target detection; Send the corresponding original water surface images to be enhanced in the training set into the generator to generate restored images; Freeze the network parameters of the multi-aspect ratio detector, and transmit the restored images and clear images in the training set to the multi-scale enhancement discriminator at the same time. Obtain the feature maps of multiple scales of the restored images and clear images through the shared backbone network, calculate the error between the two and transmit it to the generator for the generator to update the network parameters; Repeat to update the network parameters of the generator, and when the best network training parameters are obtained, the network training is completed.
[0036] In the above technical solution, the specific test steps of the network model of the general integrated water surface image enhancement and target detection system and device are as follows: Send the original water surface images to be enhanced in the test set into the generator to generate the required restored images; Input the generated restored images into the multi-aspect ratio detector network to output the target detection and recognition results.
[0037] Among them, the training of the network uses the Adam optimizer. The training of the head network lasts for 100 epochs, the Batch size is 4, and the training of the remaining parameters lasts for 150 epochs, the Batch size is 1. In the first 100 epochs, the adversarial loss, detection feature loss, and pixel mean square error loss are used to optimize the remaining parameters of the network. In the last 50 epochs, the weight of the adversarial loss is set to 0, and the detection feature loss and pixel mean square error loss are used to fine-tune the remaining network parameters. The initial learning rate during the training process is 0.001, and as the number of training epochs increases, the learning rate gradually decays at a ratio of 0.92.
[0038] In the above technical solution, collect the clear water surface images and corresponding original water surface images to be enhanced at each water surface position. Among them, the clear water surface images at the same water surface position at different times belong to different clear water surface images, and the clear water surface images at the same water surface position at different times and the corresponding original water surface images to be enhanced belong to the collection range.
[0039] In the above technical solution, such as Figure 3As shown, the generator network model and the multi-scale enhanced discriminator network model are constructed based on the generative adversarial network. The generator network model is responsible for enhancing the original water surface image; the multi-aspect ratio detector network model is responsible for detecting the targets in the initial restored image; the multi-scale enhanced discriminator network model discriminates the quality of the restored image based on the backbone network of the multi-aspect ratio detector, so as to realize the integration of water surface image enhancement and target detection.
[0040] In the above technical solution, the specific method for generating the initial restored image of each water surface position is as follows: The image to be enhanced is input into the generator network model, and the preliminary features of the image to be enhanced are obtained through a feature extraction structure composed of multiple residual convolutional blocks and instance normalization layers. The feature map is reconstructed to the original resolution of the image through a network structure composed of multiple deconvolution modules and feature extraction structures at the back end, and the initial restored image is obtained.
[0041] In the above technical solution, an efficient generator network is built through multiple feature extraction structures and deconvolution structures, which is used to efficiently construct high-quality restored images. Among them, a receptive field enhancement module is adopted in the middle, which enhances the image feature extraction ability of the generator network model and is beneficial to improving the ability of the generator to restore images; a global skip connection structure is set from the output of the front-end feature map to the middle receptive field enhancement module. The generator only needs to learn the residual between the original water surface image to be enhanced and the clear image, which can accelerate the training convergence and improve the restoration efficiency of the generator.
[0042] In the above technical solution, as Figure 4 shown, the generator network model is reconstructed on the basis of the generative adversarial network model to enhance the original water surface image. The front end of the network adopts a feature extraction structure composed of 3 residual convolutional blocks and instance normalization layers, which is responsible for extracting the preliminary features of the original water surface image to be enhanced; the middle part of the network adopts 9 receptive field enhancement modules, which are used to enhance the image feature extraction ability of the generator network model. At the same time, a global skip connection structure is set from the output of the front-end feature map to the middle receptive field enhancement module, so that the generator network model can learn the residual between the original water surface image to be enhanced and the clear image; the back end of the network uses a network structure composed of 2 deconvolution modules and feature extraction structures for image reconstruction, which is responsible for reconstructing the restored image to the original resolution to obtain the initial restored image.
[0043] In the above technical solution, the specific method for making the generator network model learn the residual between the original water surface image to be enhanced and the clear image is as follows: Take the three feature maps output by the last three branches of the backbone network of the restored image and the clear image for error calculation, and the obtained error will be transmitted to the generator to guide the generator to learn. Among them, the generator network is mainly trained by minimizing the feature loss, adversarial loss and pixel mean square error loss.
[0044] In the above technical solution, the receptive field enhancement module is divided into 5 branches. Each branch first uses convolution to perform dimensionality reduction on the input data. Based on the skip connection idea of the residual network, one of the branches after dimensionality reduction is directly connected to the front of the activation layer of the network to prevent the network performance from degrading due to the increase in the number of network layers. According to the network structure of Inception, the remaining branches respectively use convolution kernels of different sizes to perform convolution processing on the dimensionality-reduced features. Convolution kernels of different sizes can obtain receptive fields of different sizes, thereby enhancing the scale adaptability of the network to the input samples and capturing image information at different levels. At the ends of the branches, dilated convolution processing with different dilation rates is performed to expand the receptive field of the network without reducing the image resolution. After the dilated convolution processing, the remaining branches are concatenated in the channel dimension, then convolution processing is performed and fused with the data of another branch directly connected to the front of the activation layer of the network, and finally output through the ReLU activation layer.
[0045] In the above technical solution, through the convolution processing and dilated convolution of multiple branches, the scale adaptability of the generator network to the input samples is enhanced and image information at different levels is captured. Among them, convolution kernels of different sizes and dilated convolution can capture image features at different scales, enabling the network to have a good perception ability for the details and overall structure of the input image. Based on the skip connection of the residual network and the Inception structure, the performance degradation caused by the increase in the number of network layers is prevented, and the calculation efficiency is improved at the same time. Through the design of the branch structure, the network can better adapt to different types of water surface images and improve the robustness of the enhancement effect.
[0046] In the above technical solution, the receptive field enhancement module is divided into 5 branches. One of the branches is directly connected to the front of the activation layer of the network. The purpose is to use the skip connection idea to prevent the network performance from degrading due to the increase in the number of network layers. The purpose of using convolution kernels of different sizes and dilated convolution with different dilation rates for the remaining four branches is to enhance the scale adaptability of the network to the input samples and capture image information at different levels, enriching the image features extracted by the network.
[0047] In the above technical solution, one of the 5 branches uses the skip connection idea to prevent the network performance from degrading due to the increase in the number of network layers, and the remaining four branches use convolution kernels of different sizes and dilated convolution with different dilation rates to enhance the scale adaptability of the network to the input samples and capture image information at different levels, enriching the image features extracted by the network.
[0048] In the above technical solution, such as Figure 5 、 6As shown in the figure, the multi-scale enhanced discriminator network model and the multi-aspect ratio detector network model share a backbone network. Three branches are led out from the last three residual units of the backbone network. Each branch compresses the feature map of the residual unit using three consecutive convolutional layers and inputs them into the back-end networks of the multi-scale enhanced discriminator network and the multi-aspect ratio detector network respectively. Sharing the backbone network reduces the number of network parameters, enables the quality discrimination of images and the target detection to share the features in the backbone network, discriminates the image quality at multiple scales, and then conducts adversarial training with the generator based on this.
[0049] In the above technical solution, the backbone network is stacked by consecutive convolutional blocks and residual blocks, and the maximum pooling operation is discarded to retain more feature information of the image.
[0050] In the above technical solution, the multi-scale enhanced discriminator network model and the multi-aspect ratio detector network model share a backbone network, enabling the image enhancement task and the target detection task to be realized within a network framework, achieving integrated image enhancement and target detection, reducing the network parameters, reducing the network computing time consumption, and improving the efficiency of water surface image target detection and recognition.
[0051] In the above technical solution, both the multi-scale enhanced discriminator and the multi-aspect ratio detector perform subsequent processing based on the feature maps of the last three residual units of the backbone network, enabling the multi-scale enhanced discriminator to discriminate the quality of image enhancement at multiple scales based on the features related to target detection, and helping the generator to restore an image that is more conducive to target detection.
[0052] In the above technical solution, as Figure 5 shown in the figure, the front end of the multi-scale enhanced discriminator network model is the backbone network, which extracts the features of the image from shallow to deep through consecutive convolutional blocks and residual blocks, and discards the maximum pooling operation; three branches are led out from the last three residual units of the backbone network, and each branch compresses the feature map of the residual unit into a single-channel two-dimensional matrix using three consecutive convolutional layers. Each element in the matrix represents the discrimination result of the image quality within its receptive field, and the average value of all elements in the matrix is used to reflect the overall quality of the image.
[0053] In the above technical solution, by constructing a multi-scale enhanced discriminator network, the quality of image enhancement is discriminated from multiple scales. The backbone network discards the maximum pooling operation, retains more feature information of the image, reduces information loss, and is more conducive to improving the accuracy of target detection and recognition; the multi-scale enhanced discriminator network discriminates the image quality based on the feature maps of the last three residual units of the backbone network. The positions of the last 3 residual units in the backbone network are different, with different sizes of receptive fields, enabling the multi-scale enhanced discriminator to discriminate the quality of enhancement at multiple scales based on the features related to target detection, so as to help the generator restore a higher-quality image.
[0054] In the above technical solution, continuous convolutional blocks and residual blocks are used to extract features of the image from shallow to deep layers, and the maximum pooling operation is discarded, so that the multi-scale enhanced discriminator network model can identify the quality of the restored image at multiple scales based on the features related to object detection, which can help the generator restore an image that is more conducive to object detection.
[0055] In the above technical solution, as Figure 6 shown, the front end of the multi-aspect ratio detector network model is a shared backbone network for feature extraction; there are three branches in the back end. In each branch, the features are preliminarily processed by a set number of convolutional blocks. The processed deep image features are compressed and upsampled by convolutional blocks and then spliced with shallow features. At the same time, the three branches are respectively used to process water surface ship targets with different aspect ratios. Among them, the convolutional kernel scale of the branch for processing deep features is set to 1×3; the convolutional kernel scale of the branch for processing shallow features is set to 3×5; the last branch uses a convolutional kernel size of 3×3 to process the detection of square ship targets. The above network structure reduces the overall network parameters and makes the object detection more targeted.
[0056] In the above technical solution, by constructing a multi-aspect ratio detector network to detect and identify the collected images, the deep image features are compressed and upsampled by convolutional blocks and then spliced with shallow features, enriching the semantic information of the shallow features and improving the detection ability of the multi-aspect ratio detector network for targets of different scales; according to the shape characteristics of the ship targets, the detector network structure is designed specifically to process water surface ship targets with different aspect ratios, which not only reduces the detector network parameters and improves the real-time performance of the network, but also makes the object detection task more targeted, improves the adaptability of the detector to targets with different aspect ratios, and is more conducive to improving the object detection and recognition accuracy.
[0057] In the above technical solution, the discrimination result is fed back to the generator for updating the network parameters, which can improve the image restoration ability of the generator, obtain a clearer restored image, and realize the image enhancement of the system of the present invention; the object detection and recognition result is fed back to the generator for updating the network parameters, so that the generator restores an image that is more conducive to object detection and improves the object detection performance of the system of the present invention.
[0058] In the above technical solution, the multi-aspect ratio detector network is constructed based on the YOLOv3 object detection network model. The front end is the backbone network responsible for object extraction, the middle part performs feature fusion operations on the extracted feature maps, and the back end is used for classifying and positioning the objects.
[0059] Embodiment 2
[0060] A general integrated water surface image enhancement and object detection method, asFigure 2 As shown in Figure 2 , the general integrated water surface image enhancement and target detection system and device include a generator network for image restoration, a multi-scale enhancement discriminator network, and a multi-aspect ratio detector network; one or more original water surface images with motion blur collected by a camera are input into the generator network to generate an initial restored image; the generated restored image and the clear water surface images at each water surface position are input into the multi-scale enhancement discriminator network to determine whether the restored image is clear, and the discrimination result is fed back to the generator network; the generated restored image and the clear water surface images at each water surface position are synchronously input into the multi-aspect ratio detector network to output the target detection and recognition results, and the detection results are fed back to the generator model.
[0061] The general integrated water surface image enhancement and target detection method includes the following steps:
[0062] Collect clear water surface images and corresponding original water surface images to be enhanced at each water surface position;
[0063] Input the original water surface images to be enhanced at each water surface position into the generator network model to generate initial restored images at each water surface position;
[0064] Input the initial restored images at each water surface position and the corresponding clear images into the multi-scale enhancement discriminator network model to perform clarity discrimination on the initial restored images at each water surface position, and update the model parameters in the generator network model according to the clarity discrimination results to obtain a water surface image enhancement generator network model;
[0065] Input the initial restored images at each water surface position and the corresponding clear images into the multi-aspect ratio detector network model to obtain the target detection and recognition results, and the generator network model updates the model parameters in the water surface image enhancement generator network model according to the target detection and recognition results to obtain an integrated generator network model that is beneficial to water surface image enhancement and target detection of the restored image.
[0066] Embodiment 3
[0067] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in Embodiment 2.
[0068] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
Claims
1. A universal integrated water surface image enhancement and target detection system, characterized in that: It includes: The image acquisition module is used to acquire clear water surface images at various water surface positions and the corresponding original water surface images to be enhanced; The image restoration module is used to input the original water surface image to be enhanced at each water surface position into the generator network model to generate the initial restored image at each water surface position; The image discrimination module is used to input the initial restored image and the corresponding clear image of each water surface position into the multi-scale enhancement discriminator network model, and is used to perform clarity discrimination of the initial restored image of each water surface position, and update the model parameters in the generator network model according to the clarity discrimination result to obtain the water surface image enhancement generator network model; The detection and recognition module inputs the initial restored image of each water surface position and the corresponding clear image into the multi-aspect ratio detector network model to obtain the target detection and recognition results. The generator network model updates the model parameters in the water surface image enhancement generator network model according to the target detection and recognition results to obtain an integrated generator network model that is beneficial to water surface image enhancement and restored image target detection.
2. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The generator network model and the multi-scale enhanced discriminator network model are constructed based on a generative adversarial network, and the generator network model is responsible for enhancing the original water surface image; The multi-aspect ratio detector network model is responsible for detecting objects in the initial restored image; The multi-scale enhancement discriminator network model identifies the quality of the restored image based on the backbone network of the multi-aspect ratio detector, thereby realizing the integration of water surface image enhancement and target detection.
3. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The specific method of generating the initial restored image of each water surface position is as follows: the image to be enhanced is input into the generator network model, and the preliminary features of the image to be enhanced are obtained through a feature extraction structure composed of multiple residual convolution blocks and instance normalization layers. The feature map is reconstructed to the original resolution of the image through a back-end network structure composed of multiple deconvolution modules and feature extraction structures to obtain the initial restored image.
4. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The generator network model is reconstructed on the basis of the generative adversarial network model to enhance the original water surface image. The front end of the network adopts a feature extraction structure composed of multiple residual convolution blocks and instance normalization layers, which is responsible for extracting the preliminary features of the original water surface image to be enhanced. Multiple receptive field enhancement modules are used in the middle of the network to enhance the image feature extraction capability of the generator network model. At the same time, a global jump connection structure is set from the front-end feature map output to the middle receptive field enhancement module, allowing the generator network model to learn the residual between the original water surface image to be enhanced and the clear image; The network backend uses a network structure composed of multiple deconvolution modules and feature extraction structures for image reconstruction, which is responsible for reconstructing the restored image to the original resolution.
5. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The receptive field enhancement module is divided into multiple branches, each of which first uses convolution to reduce the dimension of the input data. Based on the skip connection idea of the residual network, one of the branches after the dimension reduction is directly connected to the activation layer of the network. According to the network structure of Inception, the remaining branches use convolution kernels of different sizes to perform convolution on the features after dimension reduction. Convolution kernels of different sizes can obtain receptive fields of different sizes, thereby enhancing the scale adaptability of the network to the input samples and capturing image information at different levels. At the end of each branch, dilated convolutions with different numbers of holes are performed to expand the network receptive field without reducing the image resolution. After dilated convolutions, the remaining branches are spliced in the channel dimension, and then convolved again and fused with the data of another branch directly connected to the network activation layer, and finally output through the ReLU activation layer.
6. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The multi-scale enhancement discriminator network model and the multi-aspect ratio detector network model share a backbone network, and three branches are derived from the last three residual units of the backbone network. Each branch uses three consecutive convolutional layers to compress the feature map of the residual unit, and inputs them into the back-end networks of the multi-scale enhancement discriminator network and the multi-aspect ratio detector network respectively. The shared backbone network reduces the amount of network parameters, allows image quality identification and target detection to share the features in the backbone network, performs image quality identification at multiple scales, and then conducts adversarial training with the generator based on this.
7. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The network front end of the multi-scale enhanced discriminator network model is a backbone network, which extracts features from shallow to deep layers of the image through continuous convolution blocks and residual blocks, and discards the maximum pooling operation; three branches are derived from the last three residual units of the backbone network, and each branch uses three continuous convolution layers to compress the feature map of the residual unit into a single-channel two-dimensional matrix. Each element in the matrix represents the identification result of the image quality within its receptive field, and the average value of all elements of the matrix is used to reflect the overall quality of the image.
8. The universal integrated water surface image enhancement and target detection system according to claim 1, characterized in that: The front end of the multi-aspect ratio detector network model is a shared backbone network for feature extraction; the back end has three branches, each of which uses a set number of convolution blocks to perform preliminary processing on the features, and the processed deep-level image features are compressed and up-sampled by the convolution blocks and then spliced with the shallow-level features. At the same time, the three branches are used to process surface ship targets with different aspect ratios, among which the convolution kernel scale of the branch processing deep features is set to 1×3; the convolution kernel scale of the branch processing shallow features is set to 3×5; the last branch uses a convolution kernel size of 3×3 to process the detection of square ship targets.
9. A universal integrated water surface image enhancement and target detection method, characterized in that: It includes the following steps: Collect clear water surface images at each water surface position and the corresponding original water surface images to be enhanced; The original water surface images to be enhanced at various water surface positions are input into the generator network model to generate initial restored images at various water surface positions; Inputting the initial restored image and the corresponding clear image at each water surface position into the multi-scale enhancement discriminator network model to perform clarity discrimination of the initial restored image at each water surface position, and updating the model parameters in the generator network model according to the clarity discrimination result to obtain a water surface image enhancement generator network model; The initial restored images and the corresponding clear images of each water surface position are input into the multi-aspect ratio detector network model to obtain the target detection and recognition results. The generator network model updates the model parameters in the water surface image enhancement generator network model according to the target detection and recognition results, and obtains an integrated generator network model which is beneficial to water surface image enhancement and restored image target detection.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in claim 9 are implemented.