Image enhancement device based on bionic vision, underwater target detection system and method

By simulating an image enhancement device of advanced biovision system, the accuracy and computational complexity problems of the underwater object detection model in complex environments and small object detection are solved, and efficient real-time detection in underwater environments is achieved.

CN120355593AActive Publication Date: 2025-07-22XIAMEN UNIV

Patent Information

Application Number
CN202510849219.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When facing complex environments, small underwater targets and similarity between targets and backgrounds, the existing underwater targets have low detection accuracy and high computational complexity, making it difficult to run on edge devices in real time.

Method used

The image enhancement device based on bionic vision is adopted, including the area perception module, the multi-scale feature aggregation module, the convolutional block attention mechanism module and the contrast adaptive normalization module, to simulate the functions of the advanced biological vision system, dynamically adjust the image processing to improve detection accuracy and reduce the computational complexity.

Benefits of technology

While improving target detection accuracy, it reduces computational complexity, making it suitable for real-time applications in underwater environments, especially on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355593A_ABST
    Figure CN120355593A_ABST
Patent Text Reader

Abstract

The invention discloses an image enhancement device based on bionic vision, and an underwater target detection system and method, which are used for simulating a related mechanism for extracting key information from a complex visual scene by an advanced visual animal by simulating a visual perception mechanism of an advanced visual organism, and simulating that the key information is extracted from the complex visual scene when the image enhancement device processes different visual information. Different attention characteristics are given according to the importance and the position of a target, so that the problem that an existing underwater target detection model is difficult to detect in a complex environment, underwater small targets and target and background similarity is solved. According to the method, the target detection precision is improved, meanwhile, too much calculation complexity is not brought, and the method is more suitable for real-time application in the underwater environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater target detection, and particularly relates to an image enhancement device, an underwater target detection system and a method based on bionic vision. Background Art

[0002] Underwater target detection is an important application of computer vision in underwater environments, and is widely used in fields such as marine resource exploration, ecological environment monitoring, underwater robot navigation, and marine rescue. However, the complexity of the underwater environment poses many challenges to target detection, including uneven illumination, color attenuation, low contrast, etc. In addition, underwater targets are usually small in size and are easily interfered by complex backgrounds, making feature extraction and target recognition more difficult. Therefore, how to design an efficient target detection feature extraction module in the complex underwater environment is one of the key issues to improve the detection performance.

[0003] In recent years, deep learning technologies, especially target detection algorithms based on convolutional neural networks (CNNs), such as YOLO (You Only Look Once), Faster R-CNN, and SSD (Single Shot MultiBox Detector), have made certain progress in underwater target detection tasks. However, these networks are mainly designed for natural scenes, and there are still the following problems when directly applied to underwater environments: the light attenuation and scattering in the underwater environment will reduce the contrast of the image, making it difficult for traditional convolutional layers to extract effective features and affecting the detection accuracy; underwater targets are usually small, such as fish, corals, plankton, etc., while traditional detection networks are mainly optimized for larger targets and lack the ability to extract effective features for small targets, resulting in small targets being easily ignored; the underwater scene contains a large number of background elements similar to the target, such as seaweed, rocks, plankton, etc., and these elements are similar to the target in terms of color, shape, etc., which easily leads to false detections; existing detection networks usually have a large amount of computation and are not suitable for edge devices with limited computing power (such as underwater drones, embedded platforms). How to optimize the computing efficiency while maintaining the detection accuracy has become an important research direction.

[0004] To address the challenges of underwater target detection, existing research mainly explores solutions from the following methods: (1) Underwater image enhancement. Due to the problems of uneven illumination, color attenuation, and low contrast in the underwater environment, many studies focus on image preprocessing and enhancement techniques. For example, physics-based methods (such as the DCP dehazing algorithm based on the light propagation model) attempt to restore the true color, but their generalization ability to different water environments is weak; data-driven methods (such as GAN-based image enhancement) improve the visual quality by learning end-to-end mapping, but may introduce artifacts and affect the stability of detection. In addition, the brightness correction method integrating the Retinex theory can enhance the local contrast, but it is still difficult to adapt to complex underwater optical conditions. (2) Feature extraction. Many studies use high-resolution feature maps to retain detailed information or adopt multi-branch convolutions (such as Inception-like structures) to simultaneously extract local features at different scales. However, these methods usually increase the computational complexity and may lead to information redundancy, affecting the detection efficiency. Using architectures with strong global perception capabilities such as Transformer (such as Swin-Transformer) can capture long-range dependencies, but their computational complexity is high and it is difficult to run in real time on edge devices. (3) Network structure optimization. In recent years, researchers have proposed a series of solutions to optimize the network structure to improve the robustness of underwater target detection. For example, optimization methods based on the attention mechanism (such as SE, CBAM, ECA, etc.) can enhance the attention to key regions and improve the target recognition ability. However, traditional attention mechanisms are mainly designed based on natural scenes and may not be able to well separate targets from the background in the underwater environment. In addition, researchers have also tried to introduce improved backbone networks (such as EfficientNet, MobileNetV3) to improve the feature extraction ability of the model while reducing the computational cost, but these networks are still difficult to solve the problem of small target detection. Summary of the Invention

[0005] Aiming at the problems of the existing underwater target detection models facing difficulties in detecting complex environments, underwater small targets, and the similarity between targets and backgrounds, the purpose of the present invention is to provide an image enhancement device, an underwater target detection system, and a method based on bionic vision to improve the underwater target detection ability.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is: An image enhancement device based on bionic vision, which includes a region perception module, a multi-scale feature aggregation module, a convolutional block attention mechanism module, and a contrast adaptive normalization module; The area perception module: It mimics the functions of the central retina and the peripheral retina in the advanced biological visual system. The central retina is responsible for the perception of fine target details, while the peripheral retina is used to perceive the overall scene and quickly locate the target; The multi-scale feature aggregation module: It simulates the ability of the biological visual system to process information at different scales and depths, helps to extract and aggregate features at multiple scales, so as to comprehensively capture image details; The convolutional block attention mechanism module: It simulates the attention mechanism of the advanced biological visual system, can dynamically focus on important features, suppress irrelevant information, and enhance the expression ability of key features through channel and spatial attention; The contrast adaptive normalization module: It simulates the mechanism of adaptively adjusting contrast in the biological visual system, automatically adjusts the image brightness and contrast according to different lighting conditions, and enhances the visibility of important features; The area perception module performs enhancement processing on the input feature map to obtain an enhanced feature map; the multi-scale feature aggregation module processes the enhanced feature map, fuses local and global features through different receptive field features to obtain fused features; the convolutional block attention mechanism module further screens the fused features to screen out important feature information; the contrast adaptive normalization module performs enhancement processing on the screened important feature information to obtain the final enhanced features.

[0007] The processing of the area perception module for the image is as follows: Assume the input feature map has a size of , where is the height, is the width, is the number of channels; first, use a 3×3 convolution to simulate the central retina's focusing on target details and extract local detail features, that is, ; Among them, represents a 3×3 convolution, represents background features, represents the input feature map; Next, use a 5×5 convolution to simulate the peripheral retina's perception of the global background and extract background features in a larger range:

[0008] Among them, represents a 5×5 convolution, represents background features, represents the input feature map; generate an attention map through the Sigmoid activation function: ; Among them, represents an activation function, is the local feature, is the background feature; Perform element-wise multiplication on the input feature map and the attention map to enhance the target region features and suppress the background region features, obtaining an attention-weighted feature map:

[0009] Among them, is the attention-weighted feature map, is the input feature map, is the attention map.

[0010] The region perception module uses depthwise separable convolution to replace traditional convolution, and the 5×5 convolution is replaced by a dilated convolution with a dilation rate of 2 and a kernel size of 3×3.

[0011] The processing of the multi-scale feature aggregation module is as follows: First, use three depthwise separable convolutions with dilation rates of 1, 2, and 3 respectively to extract feature information with different receptive fields: ; ; ; Among them, respectively represent the feature information of 3 different scales extracted, represents depthwise separable convolution, is the attention-weighted feature map, represents the dilation rate; Next, perform global average pooling operation to obtain the global information of the input feature map; ; Among them, represents global average pooling; represents the global information of the input feature map, is the attention-weighted feature map; Then, transform the pooling result into a one-dimensional feature representation and calculate the weights of the three scales through a fully connected layer:

[0012] Among them, represents the weights of the three scales, represents the fully connected layer, represents the global information of the input feature map, represents the activation function; Using the calculated weights Perform weighted fusion on multi-scale features:

[0013] Among them, is the weighted feature map; is the feature weight extracted when the dilation rate is ; represents the feature information extracted when the dilation rate is ; The weighted feature map will undergo further feature fusion through a 1×1 convolution, and the fused feature map will be connected residually with the input feature map to avoid information loss:

[0014] Among them, is the fused feature map, is the weighted feature map, is the attention-weighted feature map, represents a 1×1 convolution.

[0015] The convolutional block attention mechanism module includes a channel attention unit and a spatial attention unit; The main objective of the channel attention unit is to weight each channel of the feature map according to the importance of each channel, and the weight of each channel reflects the importance of that channel in the current task; The objective of the spatial attention unit is to dynamically focus on the important regions in the feature map by introducing an attention mechanism in the spatial dimension.

[0016] The specific implementation steps of the channel attention unit are as follows: Global pooling: First, the input feature map generates a global description through global average pooling and global max pooling operations; global average pooling calculates the average value of each channel to capture the global information of the image; while global max pooling calculates the maximum value of each channel to capture the most prominent features; Fully connected layer: Next, these two pooling results are input into a two-layer fully connected network; the first layer reduces the number of channels, and the second layer restores the number of channels. Each channel obtains a weight representing its importance after passing through this network; Activation function: Finally, through the Sigmoid activation function, a channel-level attention weight coefficient is obtained; this weight coefficient is multiplied with the input feature map channel by channel to obtain the adjusted feature map;

[0017] Among them, represents channel attention, , Represent the weight matrices of the two fully connected layers, represents the sigmoid activation function, represents the relu activation function, represents global average pooling, represents global maximum pooling, Fusion feature map.

[0018] The specific implementation of the spatial attention module is as follows: Channel operation of feature map: By operating the input feature map in the channel dimension, the importance of each position is calculated; the global information is obtained by average pooling in the channel dimension, and the significant features are obtained by maximum pooling in the channel dimension; Convolution operation: connect the feature maps of the two channels, extract the spatial relationship in a larger range through a 7×7 convolution, generate a spatial attention map, and output an attention map with the same spatial dimension as the input feature map, indicating the importance of each spatial position; Activation function: Finally, the sigmoid activation function is used to map the value of the spatial attention map to the interval [0, 1] and perform pixel-by-pixel multiplication with the spatial dimension of the input feature map to enhance the features of important areas and suppress unimportant areas.

[0019] in, represents channel attention, Indicates spatial attention, represents the attention of the i-th channel, Indicates the number of channels, represents the sigmoid activation function, Represents convolution.

[0020] The contrast adaptive normalization module obtains the brightness information of each channel through global mean pooling, and then adaptively adjusts the image based on this brightness information; specifically, each pixel value in the image is differentially adjusted from the global mean of its channel, and the adjusted image brightness is adapted to different input scenes by scaling the scale and bias parameters; ; in, Represents the final output, represents spatial attention, represents global average pooling, represents the zoom scale, Represents the bias parameter.

[0021] An underwater target detection system based on bionic vision, which includes a target detection model. The target detection model includes a backbone network, a neck network, and a head network. The image enhancement device described above is added to the neck network.

[0022] An underwater target detection method based on bionic vision. The method is implemented by using the underwater target detection system based on bionic vision described above. Specifically, the method is as follows: Input the original underwater image into the target detection model. The backbone network of the target detection model extracts the basic features of the shallow layer and the deep layer in the original underwater image, and then sends them to the neck network for feature fusion. The features of the image are strengthened through the image enhancement device. Finally, the strengthened features are input into the head network for classification and localization to perform target detection.

[0023] After adopting the above solution, the present invention simulates the visual perception mechanism of advanced visual organisms to simulate the relevant mechanisms of advanced visual animals for extracting key information from complex visual scenes, and bionics its characteristic of often giving different attentions according to the importance and position of the target when processing different visual information, so as to improve the problems of the existing underwater target detection model facing difficulties in detecting complex environments, small underwater targets, and the similarity between the target and the background. While improving the target detection accuracy, the present invention does not bring too much computational complexity, making it more suitable for real-time applications in underwater environments. Brief Description of the Drawings

[0024] Figure 1 It is the implementation flowchart of the system of the present invention; Figure 2 It is the structural diagram of the underwater target detection system; Figure 3 It is the flowchart of the underwater target detection; Figure 4 It is the comparison chart of the experimental result indexes; Figure 5 It is the comparison chart of the experimental effects (left: YOLOv11n, right: YOLOv11n + BVM); Figure 6 It is the comparison of the output heat maps before and after adding BVM. Detailed Embodiments

[0025] Such as Figure 1As shown in the figure, the present invention discloses an image enhancement device based on bionic vision, which includes a regional perception module, a multi-scale feature aggregation module, a convolutional block attention mechanism module, and a contrast adaptive normalization module. The regional perception module generates an enhanced feature map from the input feature map and sends it to the multi-scale feature aggregation module to further fuse local and global features through features with different receptive fields, and sends the fused features to the convolutional block attention mechanism module to further screen important feature information, and finally enters the contrast adaptive normalization module for final feature enhancement to assist the detector in obtaining clearer target features.

[0026] Among them, the regional perception module (Regional Perception Module, RPM): mimics the functions of the central retina and the peripheral retina in the advanced biological vision system. The central retina is responsible for fine target detail perception, while the peripheral retina is used to perceive the overall scene and quickly locate the target. The multi-scale feature aggregation module (Multi-Scale Feature Aggregation, MSFA): simulates the ability of the biological vision system to process information at different scales and depths, helps to extract and aggregate features at multiple scales, so as to comprehensively capture image details. The convolutional block attention mechanism module (Convloutional Block Attention Module, CBAM): simulates the attention mechanism of the advanced biological vision system, can dynamically focus on important features, suppress irrelevant information, and enhance the expression ability of key features through channel and spatial attention. The contrast adaptive normalization module (Contrast Adaptive Normalization Module, CANM): simulates the mechanism of adaptively adjusting contrast in the biological vision system, automatically adjusts the image brightness and contrast according to different lighting conditions, and enhances the visibility of important features.

[0027] The low contrast and background complexity of underwater images make it necessary for the vision system to be able to dynamically adjust the focus in the image to enhance the detail perception in low-contrast areas. There are many small targets in the underwater environment and the difference between the target and the background is small. Therefore, it is crucial to focus on the features of the target area.

[0028] In the advanced biological vision system, the fovea retina is responsible for high-resolution fine visual perception, mainly used to capture the detailed information of the target. The peripheral retina is used to perceive the overall scene, which can quickly locate the target but has a lower resolution. This division of labor enables advanced visual animals to efficiently balance local details and global perception during the visual cognition process, improving the accuracy of target detection. The region perception module precisely borrows this mechanism to help the network automatically focus on small targets or key regions in complex backgrounds, enhancing the detection ability of small underwater targets.

[0029] Specifically, the region perception module processes the image as follows: Assume the input feature map has a size of , where is the height, is the width, and is the number of channels. First, use a 3×3 convolution (fovea retina detail focusing) to extract local features :

[0030] where, represents the 3×3 convolution, represents the background feature, and represents the input feature map.

[0031] Then use a 5×5 convolution (peripheral retina background enhancement) to extract a larger range of background features :

[0032] where, represents the 5×5 convolution, represents the background feature, and represents the input feature map.

[0033] Generate an attention map through the Sigmoid activation function :

[0034] where, represents the activation function, is the local feature, and is the background feature.

[0035] Multiply the input feature map and the attention map element-wise to enhance the target region features and suppress the background region features, obtaining an attention-weighted feature map:

[0036] where, is the attention-weighted feature map, is the input feature map, is the attention map.

[0037] Meanwhile, in order to make the module more lightweight, depthwise separable convolutions are used to replace traditional convolutions, and dilated convolutions with a dilation rate of 2 and a size of 3×3 are used to replace 5×5 convolutions, so as to reduce the computational consumption.

[0038] Underwater images are often affected by light attenuation, resulting in the image being dim and blurred in some areas, making it difficult to identify image details. At the same time, there is a problem of large scale variation of underwater objects. Therefore, multi-scale feature aggregation can help extract and aggregate information at different scales, capture different details, and reduce the interference of background noise. The biological visual system is a very complex and highly optimized perception system, and one of its important characteristics is the ability to process information at different levels and scales. The multi-scale feature aggregation module is designed based on this, and by introducing different dilated convolutions and global pooling to simulate the biological visual system's processing ability for multi-scale information.

[0039] The multi-scale feature aggregation module first uses three depthwise separable convolutions with dilation rates of 1, 2, and 3 respectively to extract multi-scale information, and the input is the attention-weighted feature map output by the region perception module :

[0040]

[0041]

[0042] where respectively represent the three different scales of feature information extracted, represents the depthwise separable convolution, represents the dilation rate.

[0043] Next, a global average pooling operation is performed to obtain the global information of the input feature map.

[0044]

[0045] where, represents the global average pooling; represents the global information of the input feature map, is the attention-weighted feature map.

[0046] Then, the pooling result is transformed into a one-dimensional feature representation, and mapped to a 3D weight space through a fully connected layer to calculate the weights of the three scales:

[0047] Among them, represents the weights of three scales, represents the fully connected layer, represents the global information of the input feature map, represents the activation function.

[0048] Using the calculated weights to perform weighted fusion on the multi-scale features: .

[0049] Among them, is the weighted feature map; is the feature weight extracted when the dilation rate is , represents the feature information extracted when the dilation rate is .

[0050] The weighted feature map will undergo further feature fusion through a 1×1 convolution, and perform a residual connection between the fused feature map and the input feature map to avoid information loss: .

[0051] Among them, is the final fused feature, is the weighted feature map, is the attention-weighted feature map, represents the 1×1 convolution.

[0052] Underwater targets are often interfered by the surrounding complex background. In particular, background elements such as seaweed and plankton in the image may be very similar to the color and shape of the target. Therefore, in order to more effectively extract and identify underwater targets, a convolutional block attention mechanism module is further introduced. It enhances the network's attention to important features by introducing channel attention and spatial attention, while suppressing irrelevant or redundant features, thereby improving the expressive ability of the feature map and enhancing the accuracy of the underwater target detection network.

[0053] The convolutional block attention mechanism module includes a channel attention unit and a spatial attention unit.

[0054] Among them, the main goal of the channel attention unit is to weight each channel of the feature map according to the importance of each channel. The weight of each channel reflects the importance of that channel in the current task. The specific implementation steps are as follows: Global pooling: First, the input feature map generates global descriptions through global average pooling and global max pooling operations. Global average pooling calculates the average value of each channel to capture the global information of the image; while global max pooling calculates the maximum value of each channel to capture the most prominent features.

[0055] Fully connected layer: Next, these two pooling results are input into a two-layer fully connected network. The first layer reduces the number of channels, and the second layer restores the number of channels. Each channel obtains a weight representing its importance after passing through this network.

[0056] Activation function: Finally, through the Sigmoid activation function, a channel-level attention weight coefficient is obtained. This weight coefficient is multiplied with the input feature map channel by channel, thereby obtaining the adjusted feature map.

[0057] The purpose of channel attention is to enable the network to focus on more important channels and suppress unimportant channels, thereby improving the feature expression ability.

[0058]

[0059] Among them, represents channel attention, 、 respectively represent the weight matrices of the two fully connected layers, represents the sigmoid activation function, represents the relu activation function, represents global average pooling, represents global max pooling, Fused feature map.

[0060] The goal of the spatial attention unit is to dynamically focus on important regions in the feature map by introducing an attention mechanism in the spatial dimension. The specific implementation is as follows: Channel operations on the feature map: By operating on the input feature map in the channel dimension, the importance of each position is calculated. Global information is obtained through average pooling on the channel dimension, and significant features are obtained through max pooling on the channel dimension.

[0061] Convolution operation: These two channel feature maps are concatenated, and a 7×7 convolution is used to extract spatial relationships in a larger range to generate a spatial attention map, outputting an attention map with the same spatial dimension as the input feature map, representing the importance of each spatial position.

[0062] Activation function: Finally, through the Sigmoid activation function, the values of the spatial attention map are mapped to the [0,1] interval and multiplied with the spatial dimension of the input feature map pixel by pixel to enhance the features of important regions and suppress unimportant regions.

[0063]

[0064] in, represents channel attention, represents spatial attention, represents the attention of the i-th channel, Indicates the number of channels, represents the sigmoid activation function, Represents convolution.

[0065] The design of the contrast adaptive normalization module aims to enhance the brightness and contrast characteristics of different areas in the image by adaptively adjusting the contrast of the image. Especially when processing underwater images, it can help reduce the effects of uneven lighting, blur, color distortion and other problems in underwater images. Its main goal is to enhance the recognizability of the target by adaptively adjusting the brightness of the image.

[0066] The contrast adaptive normalization module obtains the brightness information of each channel through global mean pooling, and then adaptively adjusts the image based on this brightness information. Specifically, each pixel value in the image will be differentially adjusted from the global mean of its channel, and the adjusted image brightness will be adapted to different input scenarios through scaling and bias parameters. Through adaptive scaling and bias parameters, this module can enhance the contrast of the image, enhance the performance of the details in the image, and make the brightness of the image in different areas adapt to different scene requirements, thereby improving the accuracy of target detection or other tasks.

[0067]

[0068] in, Represents the final output, represents the spatial attention feature map of the input, represents global average pooling, represents the zoom scale, Represents the bias parameter.

[0069] In summary, the present invention simulates the visual perception mechanism of advanced visual organisms to simulate the relevant mechanism of advanced visual animals extracting key information from complex visual scenes, and imitates their characteristics of paying different attention to the importance and position of the target when processing different visual information, thereby improving the existing underwater target detection model in the face of complex environments, small underwater targets, and the detection difficulties of the similarity between the target and the background. While improving the accuracy of target detection, the present invention does not bring too much computational complexity, making it more suitable for real-time applications in underwater environments.

[0070] The underwater target detection system includes a target detection model, and the target detection model includes a backbone network, a neck network, and a head network. When applying the image enhancement device of the present invention to underwater target detection, the image enhancement device can be added to the neck network. For example, as Figure 2 shown, the image enhancement device (BVM) is added to the neck network and is located before the head network. At the same time, there are also other combination methods. For example, it can be combined into the network architecture, such as C3k2 or other parts of the network. The specific insertion method is flexible and can be adjusted according to the actual situation. Figure 2 In the figure, Input represents the input, Conv represents convolution, C3k2 represents a lightweight feature extraction structure, C2PSA represents a lightweight feature enhancement module integrating a spatial attention mechanism, SPPF represents an efficient feature enhancement module, Concat represents feature map splicing, Upsample represents upsampling, and Detect represents the detection head.

[0071] For example, as Figure 3 shown, when performing underwater target detection, the original underwater image is input into the target detection model. The backbone network of the target detection model extracts the shallow and deep basic features in the original underwater image, and then sends them to the neck network for feature fusion. The features of the image are strengthened through the image enhancement device, and finally the strengthened features are input into the head network for classification and localization to perform target detection.

[0072] To verify the effectiveness of the module proposed by the present invention, the present invention selects to conduct experiments on the publicly available underwater dataset DUO. The dataset contains 7,782 accurately labeled images, of which 6,671 are used for training and 1,111 are used for testing. There are a total of four categories and 74,515 objects in the dataset. The numbers of sea cucumbers (Holothurian), sea urchins (Echinus), scallops (Scallop), and starfish (Starfish) are 7,887 (10.6%), 50,156 (67.3%), 1,924 (2.6%), and 14,548 (19.5%), respectively. The DUO dataset contains a series of diverse underwater images and corresponding annotation information. These images cover various underwater scenes, including coral reefs, seabed sediments, shipwrecks, etc., and at the same time contain objects of different shapes, sizes, and colors. The images in the dataset exhibit typical underwater image characteristics such as color deviation, low contrast, uneven illumination, blurring, and high noise, which bring certain challenges to accurately detecting different aquaculture organisms and at the same time largely reflect the problems faced by the detection targets in the real marine environment.

[0073] The present invention takes the advanced object detector YOLOv11n as the baseline. To solve the problems of complex underwater environment, large number of small targets, and object occlusion, a designed image enhancement device is added before the output of the small target detection head of the YOLOv11n detector to improve its detection ability for underwater targets and keep the model lightweight enough. In terms of experimental parameters, the present invention sets the number of training epochs to 300, the batch size to 16, uses Adam as the optimizer for the model, the initial learning rate of the model is 0.001, and the weight decay coefficient of the optimizer is 0.0005. In terms of experimental hardware, the present invention uses 1 NVIDIA GeForce Rtx 4090 GPU for training, the pytorch version is 1.11.0, and the python version is 3.8.

[0074] Meanwhile, to further compare the superiority of the designed module, the YOLOv11n version with the image enhancement device is compared with the detectors of each YOLO series version with the same lightweight.

[0075] The key of the present invention lies in its dynamic feature focusing mechanism. By flexibly adjusting the attention weights, it can adaptively focus on the key areas according to the different complexities of the samples, thus improving the accuracy of underwater target detection. The module effectively enhances the recognition ability for complex backgrounds and small targets through multi-level feature fusion, including regional perception, multi-scale feature aggregation, and convolutional block attention mechanism. In addition, the bionic vision module reduces the fluctuations during the training process and ensures the stability of feature learning through dynamic weight generation and adaptive smoothing optimization strategy. As Figures 4 - 6 shown, the experimental results show that the bionic vision module effectively improves the performance of the detector after adding the advanced detector and does not bring too much computational burden. Figure 4 Among them, mAP@0.5 refers to the average AP of each category when the IoU (Intersection over Union) threshold is 0.5. mAP@0.5:0.95 means that the IoU threshold is taken at 10 equally spaced IoU thresholds in the range of 0.5 to 0.95 with a step of 0.05, and the average AP of each category is calculated at different IoU thresholds. In the object detection task, the number of giga floating-point operations per second (GFLOPs) is usually used to measure the computational amount of the model.

[0076] The key points of the present invention are as follows: (1) Inspired by the biological vision mechanism, the design of the basic structure module of the underwater target detection enhancement module network is carried out. (2) The designed module is simple and efficient. While combining the attention mechanism and feature fusion and introducing contrast enhancement, it has the characteristics of lightweight and plug-and-play.

[0077] The present invention specifically aims at the problem that the underwater imaging environment is complex, and various factors such as light, water body, and impurities affect the imaging quality, resulting in low image contrast of underwater videos, color changes, and difficult feature recognition. It systematically imitates the target recognition process of the visual system of advanced organisms. First, through the regional perception module, it imitates the functions of the central retina and the peripheral retina in the visual system of advanced organisms. The central retina is responsible for the perception of fine target details, while the peripheral retina is used to perceive the overall scene and quickly locate the target, focus on the target area, and enhance the ability to extract key features. Secondly, through the multi-scale feature aggregation module, it ensures the extraction of information at multiple scales and then combines the attention mechanism module to dynamically focus on important features. Finally, a contrast adaptive normalization module is also added to enhance the contrast of underwater images and reduce the influence of light attenuation and blurring of underwater images.

[0078] Therefore, the multi-scale feature aggregation module of the present invention is significantly different from the multi-scale fusion methods proposed in the prior art in terms of structural design and perception mechanism. It is more biomimetic, adaptive, lightweight, and has the ability of global semantic understanding in the fusion strategy, and is particularly suitable for processing image scenes with complex backgrounds, uneven underwater lighting, and large target size changes. By introducing the dynamic weighting and global information perception mechanism, the detection ability for small targets and low-contrast targets is effectively improved, and these advantages have not appeared or have not been systematically integrated in the existing methods. The main purpose of the present invention is to construct a plug-and-play lightweight underwater target detection module driven by biological vision perception. It emphasizes more on generality, flexibility, and deployment efficiency in design, and can be used as a functional enhancement unit of the existing target detector to improve its detection robustness and accuracy in complex underwater scenes without significantly increasing the computational cost.

[0079] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. An image enhancement device based on bionic vision, characterized in that: It includes a region perception module, a multi-scale feature aggregation module, a convolutional block attention mechanism module, and a contrast adaptive normalization module; The region perception module: mimics the functions of the central retina and the peripheral retina in the advanced biological visual system. The central retina is responsible for fine target detail perception, while the peripheral retina is used to perceive the overall scene and quickly locate the target; The multi-scale feature aggregation module: simulates the ability of the biological visual system to process information at different scales and depths, helps extract and aggregate features at multiple scales, so as to comprehensively capture image details; The convolutional block attention mechanism module: simulates the attention mechanism of the advanced biological visual system, can dynamically focus on important features, suppress irrelevant information, and enhance the expression ability of key features through channel and spatial attention; The contrast adaptive normalization module: simulates the mechanism of adaptively adjusting contrast in the biological visual system, automatically adjusts the image brightness and contrast according to different lighting conditions, and enhances the visibility of important features; The region perception module performs enhancement processing on the input feature map to obtain an enhanced feature map; The multi-scale feature aggregation module processes the enhanced feature map, fuses local and global features through different receptive field features to obtain a fused feature; the convolutional block attention mechanism module further screens the fused feature to screen out important feature information; the contrast adaptive normalization module performs enhancement processing on the screened important feature information to obtain the final enhanced feature.

2. The image enhancement device based on bionic vision according to claim 1, wherein: The processing of the region perception module for the image is as follows: Assume the input feature map has a size of , where is the height, is the width, is the number of channels; first, use a 3×3 convolution to simulate the focusing of the central retina on the target details and extract local detail features, that is ; Among them, represents a 3×3 convolution, represents background features, represents the input feature map; Next, a 5×5 convolution is used to mimic the perception of the global background by the peripheral retina, and background features in a larger range are extracted: Among them, represents a 5×5 convolution, represents background features, represents the input feature map; an attention map is generated through the Sigmoid activation function : ; Among them, represents an activation function, is a local feature, is a background feature; Perform element-wise multiplication on the input feature map and the attention map to enhance the features of the target region and suppress the features of the background region, obtaining an attention-weighted feature map: Among them, is the attention-weighted feature map, is the input feature map, is the attention map.

3. The image enhancement device based on bionic vision according to claim 2, characterized in that: The region perception module uses depthwise separable convolution to replace traditional convolution, and the 5×5 convolution is replaced by a dilated convolution with a dilation rate of 2 and a kernel size of 3×3.

4. The image enhancement device based on bionic vision according to claim 1, characterized in that: The processing of the multi-scale feature aggregation module is specifically as follows: First, three depthwise separable convolutions with dilation rates of 1, 2, and 3 are used to extract feature information with different receptive fields: ; ; ; Among them, respectively represent the feature information of three different scales extracted, represents depthwise separable convolution, is the feature map weighted by attention, represents the dilation rate; Next, a global average pooling operation is performed to obtain the global information of the input feature map; ; Among them, represents global average pooling; represents the global information of the input feature map, is the attention-weighted feature map; Then, the pooling result is transformed into a one-dimensional feature representation, and a fully connected layer is used to calculate the weights of the three scales: Among them, represents the weights of three scales, represents the fully connected layer, represents the global information of the input feature map, represents the activation function; Using the calculated weights Perform weighted fusion on multi-scale features: Among them, is the weighted feature map; is the feature weight extracted when the dilation rate is ; represents the feature information extracted when the dilation rate is ; The weighted feature map will undergo further feature fusion through a 1×1 convolution, and the fused feature map is connected with the input feature map in a residual manner to avoid information loss: Among them, is the fused feature map, is the weighted feature map, is the attention-weighted feature map, represents a 1×1 convolution.

5. The image enhancement device based on bionic vision according to claim 1, wherein: The convolutional block attention mechanism module includes a channel attention unit and a spatial attention unit; The main goal of the channel attention unit is to weight each channel of the feature map according to the importance of each channel, and the weight of each channel reflects the importance of this channel in the current task; The goal of the spatial attention unit is to dynamically focus on important regions in the feature map by introducing an attention mechanism in the spatial dimension.

6. The image enhancement device based on bionic vision according to claim 5, wherein: The specific implementation steps of the channel attention unit are as follows: Global pooling: First, the input feature map is processed through global average pooling and global maximum pooling operations to generate a global description; global average pooling calculates the average value of each channel to capture the global information of the image; while global maximum pooling calculates the maximum value of each channel to capture the most prominent features; Fully connected layer: Next, these two pooling results are input into a two-layer fully connected network; The first layer reduces the number of channels, and the second layer restores the number of channels. After passing through this network, each channel is given a weight indicating its importance. Activation function: Finally, a channel-level attention weight coefficient is obtained through the Sigmoid activation function; this weight coefficient is multiplied channel by channel with the input feature map to obtain the adjusted feature map; Among them, represents channel attention, and represent the weight matrices of two fully connected layers respectively, represents the sigmoid activation function, represents the relu activation function, represents global average pooling, represents global max pooling, is the fused feature map.

7. An image enhancement device based on bionic vision according to claim 5, characterized in that: The specific implementation of the spatial attention module is as follows: Channel operation of feature map: By operating the input feature map in the channel dimension, the importance of each position is calculated; the global information is obtained by average pooling in the channel dimension, and the significant features are obtained by maximum pooling in the channel dimension; Convolution operation: connect the feature maps of the two channels, extract the spatial relationship in a larger range through a 7×7 convolution, generate a spatial attention map, and output an attention map with the same spatial dimension as the input feature map, indicating the importance of each spatial position; Activation function: Finally, the sigmoid activation function is used to map the value of the spatial attention map to the interval [0, 1] and perform pixel-by-pixel multiplication with the spatial dimension of the input feature map to enhance the features of important areas and suppress unimportant areas. Among them, represents channel attention, represents spatial attention, represents the attention of the i-th channel, represents the number of channels, represents the sigmoid activation function, represents convolution.

8. An image enhancement device based on bionic vision according to claim 1, characterized in that: The contrast adaptive normalization module obtains the brightness information of each channel through global mean pooling, and then adaptively adjusts the image based on this brightness information; specifically, each pixel value in the image is differentially adjusted from the global mean of its channel, and the adjusted image brightness is adapted to different input scenes by scaling the scale and bias parameters; ; Among them, represents the final output, represents spatial attention, represents global average pooling, represents the scaling factor, represents the bias parameter.

9. An underwater target detection system based on bionic vision, characterized in that: It comprises a target detection model, which comprises a backbone network, a neck network and a head network, and the neck network is added with an image enhancement device as described in any one of claims 1-8.

10. An underwater target detection method based on bionic vision, characterized in that: The method is implemented by using an underwater target detection system based on bionic vision as claimed in claim 9, and the method is specifically as follows: The original underwater image is input into the target detection model. The backbone network of the target detection model extracts the basic features of the shallow and deep layers in the original underwater image, and then sends them to the neck network for feature fusion. The image features are enhanced through the image enhancement device, and finally the enhanced features are input into the head network for classification and positioning for target detection.

Citation Information

Patent Citations

  • Eagle-eye-vision-imitating double-fovea-center air target saliency detection method

    CN116863354A

  • Image recognition method based on bionic vision and fluctuation enhancement

    CN119152354A

  • Small target detection method based on multi-scale cavity fusion

    CN119992390A

  • Multi-task joint sensing network model and detection method for traffic road surface information

    WO2024138993A1

Cited By

  • Metal surface defect detection method based on improved DETR model

    CN121438018A