Complex underwater target detection method and system based on multi-scale features
By combining relative global histogram stretching preprocessing of underwater images with a multi-scale feature extraction network, the problems of low contrast, color distortion and occlusion in underwater target detection are solved, achieving high-precision target detection with a low false negative rate.
Patent Information
- Application Number
- CN202511017765.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
AI Technical Summary
Existing underwater target detection technologies suffer from low detection accuracy and high false negative rates when faced with problems such as low contrast, color distortion, light scattering, and target occlusion, making it difficult to meet real-time requirements.
A relative global histogram stretching algorithm is used for preprocessing, combined with a multi-scale feature extraction network, including a multi-gradient interactive structure and a progressive feature fusion module, to improve image quality and enhance feature extraction and fusion capabilities.
It significantly improves the accuracy and robustness of underwater target detection, reduces the false negative rate, and meets the real-time detection needs in complex underwater environments.
Smart Images

Figure CN120912862A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of underwater target detection, and particularly relates to a complex underwater target detection method and system based on multi-scale features. BACKGROUND
[0002] In recent years, underwater target detection technology plays an increasingly important role in target detection in multiple fields. By using visual autonomous perception technology, underwater targets can be accurately identified and analyzed. This method not only effectively reduces manpower, but also reduces the risks associated with high-risk underwater operations. At present, a series of sensors (such as cameras and sonars) are usually used to collect underwater environmental information. Through this method, continuous, long-term and real-time observation can be achieved. Although sonar-based detection methods provide a wide detection range, they are relatively inaccurate. Therefore, they are not suitable for detecting small underwater targets.
[0003] For close-range underwater target detection, using optical images with rich feature information is a more effective choice. However, due to the particularity and complexity of the underwater environment and the lighting conditions. At the same time, underwater optical images often have low contrast and serious color deviation problems. This causes the details of the detected target to be blurred, and the detection accuracy is limited. In addition, the differences between underwater targets, the low distinction between targets and backgrounds, and the occlusion between targets further increase the miss rate.
[0004] Due to the differences in the propagation characteristics of light in water and air. At the same time, the particles in the water will scatter light. These will cause the collected underwater images to have poor contrast, color changes and fogging. These factors have a great impact on the subsequent target detection results. Therefore, a lot of research has been conducted to provide practical methods to meet the image enhancement requirements of a series of complex underwater scenes.
[0005] Underwater target detection algorithms based on deep learning have gradually become the focus of research, because traditional target detection algorithms have significant time complexity in feature extraction, which cannot meet the requirements of large-scale data processing and real-time. Among them, the region-based method is divided into two stages, represented by the R-CNN system. They include Fast R-CNN, Faster R-CNN, Mask R-CNN, etc. Although they can be applied to underwater target detection technology, the detection speed is slow and cannot meet the real-time requirements in application scenarios. YOLO series as a representative of one-stage target detection algorithm based on regression has been widely applied. They can be applied to underwater target detection.
[0006] Despite the significant progress made by researchers in the field of underwater target detection and image enhancement, there are still some problems to be solved. First, different wavelengths of light attenuate differently when propagating underwater, making the image appear blue-green in color and low in contrast. At the same time, different particles in the water will scatter light, causing the edges of the image to become blurred and producing a fog-like effect. It will blur the local details of the detected target and reduce the accuracy of detection. In addition, some underwater targets have different sizes. There is occlusion between them, resulting in a high false negative rate of target detection. SUMMARY
[0007] To solve the technical problems existing in the prior art, the present application proposes a complex underwater target detection method and system based on multi-scale features, aiming to enhance the network's ability to distinguish features at various scales and improve detection performance.
[0008] In one aspect, to achieve the above-mentioned purpose, the present application provides a complex underwater target detection method based on multi-scale features, comprising:
[0009] Obtain an underwater optical image to be detected, and pre-process the underwater optical image using a relative global histogram stretching algorithm to obtain a pre-processed underwater optical image;
[0010] Construct a multi-scale feature extraction network;
[0011] Input the pre-processed underwater optical image into the multi-scale feature extraction network for processing to generate a multi-scale target detection result;
[0012] The multi-scale feature extraction network includes a feature extraction module and a multi-scale feature fusion module. The feature extraction module extracts multi-scale features through a multi-gradient interaction structure, and the multi-scale feature fusion module uses a progressive fusion strategy to integrate the multi-scale features.
[0013] Preferably, the relative global histogram stretching algorithm is used to pre-process the underwater optical image, comprising:
[0014] The RGB channel of the underwater optical image is decomposed into independent channels and color balance corrected;
[0015] The corrected image channel is dynamically histogram stretched, and the noise is removed and the details are preserved using bilateral filtering;
[0016] The L component is linearly stretched in the CIE Lab color space, and the a and b components are respectively corrected by S-shaped curve, and finally converted back to the RGB model to obtain the processed underwater optical image.
[0017] Preferably, the RGB channels of the underwater optical image are decomposed into independent channels and color balance correction is performed, including:
[0018]
[0019] wherein R avg , G avg , and B avg represent the normalized average values of the recovered red, green, and blue channels respectively, M and N represent the spatial resolution of the image, I g and I b represent the gray values of green and blue of each pixel point respectively, i represents the column where the pixel is located, j represents the row where the pixel is located, and θ g is the color balance coefficient of the G channel, and θ b is the color balance coefficient of the B channel.
[0020] Preferably, the dynamic histogram stretching is performed as follows:
[0021]
[0022] wherein p in and p out represent the input pixel and the output pixel respectively, I min , I max , Q min , and Q max are adaptive parameters of the image before and after stretching.
[0023] Preferably, the linear stretching is performed on the L component in the CIE Lab color space as follows:
[0024]
[0025] wherein RD is the stretching result of the pixel value, x is the gray value, and a is the darkest gray value.
[0026] The S-shaped curve correction is performed as follows:
[0027]
[0028] wherein p x is the output pixel, I x is the input pixel, a is the darkest gray value, and b is the brightest gray value.
[0029] Preferably, the feature extraction module comprises a backbone network and a multi-gradient interaction structure, wherein the backbone network adopts a convolutional neural network to extract multi-scale features, the multi-gradient interaction structure is composed of three parallel branches, the feature interaction is realized by shuffling operation after channel splicing of the outputs of the branches, and the semantic feature weighting is performed by using a SimAM attention mechanism.
[0030] Preferably, the multi-gradient interaction structure comprises:
[0031] The first branch comprises a 1*1 convolution for extracting first-layer features;
[0032] The second branch comprises a 1*1 convolution and four 3*3 convolution blocks, and two gradient paths are introduced simultaneously, for extracting second-layer features;
[0033] The third branch comprises a 1*1 convolution and four 5*5 convolutions, and four gradient paths are introduced simultaneously, for extracting third-layer features.
[0034] Preferably, the multi-scale feature fusion module fuses the first-layer features, the second-layer features and the third-layer features in sequence through an adaptive feature pyramid network (AFPN), performs upsampling and downsampling operations on the features to unify the feature map size, and realizes cross-scale adaptive spatial fusion through an ASFF3 unit.
[0035] Preferably, in the feature fusion process, an adaptive weight distribution mechanism is adopted to dynamically adjust the contribution degree of different scale features according to the feature response intensity.
[0036] On the other hand, to achieve the above-mentioned purpose, the application further provides a complex underwater target detection system based on multi-scale features, comprising:
[0037] An image acquisition device is configured to acquire an underwater optical image to be detected, and to preprocess the underwater optical image by using a relative global histogram stretching algorithm to obtain a preprocessed underwater optical image.
[0038] A model construction device is configured to construct a multi-scale feature extraction network.
[0039] A result generation device is configured to input the preprocessed underwater optical image into the multi-scale feature extraction network for processing to generate a multi-scale target detection result.
[0040] The multi-scale feature extraction network comprises a feature extraction module and a multi-scale feature fusion module, the feature extraction module extracts multi-scale features through a multi-gradient interaction structure, and the multi-scale feature fusion module integrates the multi-scale features by using a progressive fusion strategy.
[0041] Compared with the prior art, the application has the following advantages and technical effects:
[0042] (1) The present application applies a relative global histogram stretching algorithm to the data set used in the present application. By reducing the negative effects of underwater light attenuation scattering caused by suspended particles, the image quality is improved. The relative global histogram stretching algorithm greatly improves the image contrast and reduces the color distortion, which provides higher quality input data for the subsequent target detection task. In addition, by solving the blurring effect of light scattering, the visibility of the target contour is effectively improved, which lays a solid foundation for feature extraction in network modeling;
[0043] (2) In order to further improve the performance of target detection, the present application integrates a multi-gradient interaction structure into the backbone network, which supports multi-gradient feature aggregation and helps to preserve and amplify local details in the target, especially for objects with blurred edges or complex structures. The multi-gradient interaction structure not only improves the model's attention to key semantic information during feature extraction, but also allows more effective learning of challenging detail areas, which improves the accuracy and precision of target detection;
[0044] (3) The present application optimizes the neck network structure by designing a multi-scale feature fusion module, promotes the interaction of different scale feature information, and realizes cross-channel information interaction. This multi-path feature fusion method effectively integrates low-level and high-level features, enabling the network to capture more comprehensive feature information. This improvement not only enriches the model's representation ability, but also significantly improves the detection performance of occluded targets and reduces the missed detection rate caused by occlusion. By improving the feature fusion strategy, the detection accuracy is improved, making the model more robust in complex underwater environments. BRIEF DESCRIPTION OF DRAWINGS
[0045] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The illustrative embodiments of the present application and their description serve to explain the present application. They do not, however, limit the present application, which is set forth in the claims. In the drawings:
[0046] Figure 1 A flow chart of a complex underwater environment target detection method based on artificial intelligence according to an embodiment of the present application;
[0047] Figure 2 A structure diagram of an underwater target detection algorithm based on multi-scale feature fusion according to an embodiment of the present application;
[0048] Figure 3 A flow chart of a relative global histogram stretching algorithm according to an embodiment of the present application;
[0049] Figure 4 A schematic diagram of a multi-gradient interaction structure according to an embodiment of the present application;
[0050] Figure 5 A multi-scale feature extraction network structure diagram according to an embodiment of the present application. DETAILED DESCRIPTION
[0051] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0052] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0053] Due to the light attenuation and light scattering of particles in water, underwater images face challenges such as low contrast and color distortion. These problems cause the details of the target in the image to be blurred, thereby reducing the detection accuracy. At the same time, underwater targets exist in different scales and are often occluded, resulting in an increased rate of missed detection. To address this challenge, the present embodiment proposes an artificial intelligence-based complex underwater target detection method. This method aims to enhance the network's ability to distinguish features at various scales, thereby improving detection performance. In view of the problem of color deviation and low contrast of underwater images, a relative global histogram stretching algorithm is applied to preprocess the underwater images in the dataset. Secondly, in order to solve the problem of high rate of missed detection caused by dense distribution and occlusion of underwater targets, a multi-gradient interaction structure is proposed. This method enhances the model's ability to capture complex features and guides the network to prioritize processing key semantic information. In addition, a multi-path interaction structure is also proposed, which can share multi-scale feature information generated due to the different scales of underwater targets and the blurring of detected target details. The process of realizing semantic information fusion and interaction between different scale features extracted from the backbone aims to improve the network's ability to express features.
[0054] A complex underwater target detection method based on multi-scale features, as Figure 1 , comprises:
[0055] An underwater optical image to be detected is acquired, and a relative global histogram stretching algorithm is used to preprocess the underwater optical image, obtaining a preprocessed underwater optical image;
[0056] A multi-scale feature extraction network is constructed;
[0057] The preprocessed underwater optical image is input into the multi-scale feature extraction network for processing, generating a multi-scale target detection result;
[0058] The multi-scale feature extraction network comprises a feature extraction module and a multi-scale feature fusion module. The feature extraction module extracts multi-scale features through a multi-gradient interaction structure, and the multi-scale feature fusion module integrates the multi-scale features using a progressive fusion strategy.
[0059] The structure of the algorithm proposed in this embodiment is shown in Figure 2
[0060] Further, the underwater optical image is preprocessed by using a relative global histogram stretching algorithm, including:
[0061] The RGB channel of the underwater optical image is decomposed into independent channels and color balance correction is performed;
[0062] The corrected image channel is subjected to dynamic histogram stretching, and a bilateral filter is used to eliminate noise and retain details;
[0063] The L component is linearly stretched in the CIE Lab color space, and the a and b components are respectively subjected to S-shaped curve correction, and finally converted back to the RGB model to obtain the processed underwater optical image.
[0064] Specifically, due to different degrees of attenuation caused by different wavelengths of light propagating underwater, underwater photos are usually blue-green. Light scattered by water-suspended particles can cause the underwater photos taken to have low contrast, blurred edges, and fog. Before image training, a simple and effective technique suitable for shallow water is used to improve the quality of underwater images to meet the requirements of the image quality of the target detection system. By enhancing the contrast and clarity of the image, the expression of the semantic information of the image can be enhanced, thereby ensuring that the target detection model can be successfully trained. In order to balance the requirements of image quality and model accuracy, the relative global histogram stretching image enhancement algorithm is selected in this embodiment. Figure 3 The overall workflow of the relative global histogram stretching algorithm is shown.
[0065] The relative global histogram stretching algorithm process consists of two parts, contrast correction and color correction.
[0066] In the contrast correction step, the decomposed R-G-B channels of the image are subjected to color balance. The G-B channels are corrected according to formulas (1) and (2). In this embodiment, the R channel is not considered because red light in water is difficult to compensate with simple color balance, which can cause red oversaturation. The color balance coefficient θ g 、θ b of the G-B channel is calculated using formula (3), where M and N represent the spatial resolution of the image, and Ig and Ib represent the gray values of green and blue, respectively, of each pixel point.
[0067]
[0068]
[0069] In the formula, R avg 、Gavg , B avg represent the normalized mean values of the recovered red, green and blue channels, M, N represent the spatial resolution of the image, I g and I b represent the gray value of green and blue of each pixel, i is the column where the pixel is located, j is the row where the pixel is located, θ g is the color equalization coefficient of G channel, θ b is the color equalization coefficient of B channel.
[0070] Next, dynamic histogram stretching is performed on the image channel, as formula (4):
[0071]
[0072] In the formula, p in and p out represent the input pixel and the output pixel, I min , I max , Q min and Q max are adaptive parameters of the image before and after stretching.
[0073] Subsequently, bilateral filtering is used to reduce the noise introduced by the previous transformation while preserving the basic details of the underwater color image. This technique effectively neutralizes low contrast, reduces color distortion caused by light scattering and absorption, thereby improving the overall image quality.
[0074] Further, after completing the previous steps, color correction is continued on the image. Color correction is achieved by stretching the "L" component in the CIE Lab color space and modifying the "a" and "b" components. First, the underwater image is converted to the CIELab color model, where the "L" component is used to adjust the image brightness by linear sliding stretching, L = 100 represents the brightest value, L = 0 represents the darkest value, as described in formula (5), which is a continuous probability distribution of positive random variables. Subsequently, the output color level of the "a" and "b" components is modified to achieve accurate color correction. When a = 0 and b = 0, the color channel will present a true neutral gray value. The stretching of the "a" and "b" components is defined by an S-shaped curve, as shown in formula (6).
[0075]
[0076] In the formula, RD is the stretching result of the pixel value, x is the gray value, p χ is the output pixel, I χ is the input pixel, a is the darkest gray, and b is the brightest gray.
[0077] In this embodiment, the variable is set to the optimal experimental value of 1.3. Formula (6) adopts an exponential function as the stretching coefficient; the closer the value is to 0, the better the stretching effect.
[0078] After performing adaptive stretching on each component, the channels are merged and converted back to the RGB model, and the output image obtained achieves enhanced contrast and accurate color correction.
[0079] In summary, the relative global histogram stretching enhancement algorithm shows significant advantages in underwater image processing, mainly in color restoration, contrast enhancement, and detail preservation. Due to the effects of light scattering and absorption, underwater images often suffer from color distortion, low contrast, and blurred details. By combining color equalization and histogram stretching techniques, the relative global histogram stretching algorithm effectively addresses these issues. First, color equalization helps restore the natural colors of underwater images, resulting in a more realistic representation. Second, the histogram stretching in the relative global histogram stretching algorithm enhances the brightness and contrast of the image, making initially dull underwater images appear clearer. Additionally, the algorithm preserves edge and detail information well, avoiding over-smoothing and highlighting key features in underwater scenes. Overall, the relative global histogram stretching algorithm significantly improves the visual quality of underwater images, making it particularly suitable for applications requiring high-quality underwater imaging.
[0080] Further, the feature extraction module comprises a backbone network and a multi-gradient interaction structure, wherein the backbone network adopts a convolutional neural network to extract multi-scale features, the multi-gradient interaction structure is composed of three parallel branches, feature interaction is realized by shuffling operation after channel splicing of branch outputs, and a SimAM attention mechanism is adopted for semantic feature weighting.
[0081] The multi-gradient interaction structure comprises:
[0082] The first branch comprises a 1*1 convolution for extracting first-layer features;
[0083] The second branch comprises a 1*1 convolution and four 3*3 convolution blocks, and two gradient paths are introduced simultaneously, for extracting second-layer features;
[0084] The third branch comprises a 1*1 convolution and four 5*5 convolutions, and four gradient paths are introduced simultaneously, for extracting third-layer features.
[0085] Specifically, underwater targets are densely distributed and occluded, which leads to a high rate of missed detection. To address the above problems, a multi-gradient interaction structure is designed with reference to the advantages of the Inception structure. Before concat, the structure consists of three branches. The first branch is a 1*1 convolution for extracting the first layer features. The second branch includes a 1*1 convolution and four 3*3 convolution blocks, while introducing two gradient paths on the branch for extracting the second layer features. Finally, the third branch includes a 1*1 convolution and four 5*5 convolutions for extracting the third layer features, while introducing four gradient paths to prepare for subsequent feature fusion. In this way, features of different scales can be obtained. The multi-branch structure enhances the model's ability to process complex features by promoting multi-scale feature extraction. In addition, the parameters of each convolution kernel in the network structure are shared, thereby reducing the total number of parameters learned by the model. This parameter sharing not only optimizes the efficiency of the model, but also reduces the risk of network overfitting. However, single-block feature learning is often insufficient. This embodiment combines the features extracted by the three branch paths and adds a shuffle operation to achieve feature interaction between different branches. Although the network can obtain feature information from different branches, the shuffle operation improves the information flow between feature channels. It helps the network to enhance feature representation. It also significantly reduces the computational cost while maintaining accuracy.
[0086] At the same time, multi-path feature extraction inevitably produces redundant feature information. Therefore, a SimAM attention mechanism is integrated after the shuffle operation. It realizes the weight distribution of effective semantic features extracted by different convolutions. Therefore, the network focuses on semantic features with more information and ignores ineffective features. The multi-gradient interaction structure is as shown in Figure 4
[0087] Further, the multi-scale feature fusion module fuses the first layer features, the second layer features and the third layer features in turn through the progressive feature pyramid network AFPN, performs upsampling and downsampling operations respectively to unify the feature map size, and realizes cross-scale adaptive spatial fusion through the ASFF3 unit.
[0088] During the feature fusion process, an adaptive weight distribution mechanism is adopted to dynamically adjust the contribution of different scale features according to the feature response intensity.
[0089] Specifically, in the target detection task, the extraction of multi-scale features is crucial for encoding targets with varying scales. One popular method for multi-scale feature extraction involves the use of traditional top-down and bottom-up feature pyramid networks. However, these methods often encounter challenges related to feature information loss, affecting the fusion of features between non-adjacent regions. To address this issue, the present embodiment proposes an Asymptotic Feature Pyramid Network (AFPN) that fuses two adjacent low-level features and gradually incorporates high-level features into the fusion process. To avoid significant semantic gaps between non-adjacent regions, an adaptive spatial fusion operation is introduced. This is particularly important as conflicts in multi-object information can arise during the feature fusion process at each spatial location, and the adaptive spatial fusion operation is implemented to alleviate these inconsistencies.
[0090] The AFPN network fuses the first and second layer features in the initial stage. The third layer features are fused in the subsequent stages. In the face of complex underwater environments, most optical images taken have the problem of blurred details. This will result in low detection accuracy. To solve the above problem, the neck network of YOLOv11 is improved combined with the idea of AFPN network. The specific structure is as shown in Figure 5
[0091] After extracting features from the backbone network from bottom to top, the features are output from three independent feature layers of different scales. The first and second layer features are input to the ASFF3 unit after convolution operation. Then, three different levels of convolution operation are performed. The output is adaptively spatially fused at multiple scales and adaptively fused through the ASFF3 unit. Finally, the detection targets at different scales are obtained. Generally, direct fusion between non-adjacent feature layers will reduce the subsequent detection effect. Therefore, the present embodiment uses a progressive method to fuse features of different scales. It helps to achieve effective communication between semantic information. At the same time, in order to ensure that the input and output dimensions are the same and to prepare for feature fusion, 1*1 convolution and 3*3 convolution are also used for upsampling and downsampling operations.
[0092] The present embodiment optimizes the neck network structure by designing a multi-scale feature fusion module, promotes the interaction of feature information of different scales, and realizes cross-channel information interaction. This multi-path feature fusion method effectively integrates low-level and high-level features, enabling the network to capture more comprehensive feature information. This improvement not only enriches the representation ability of the model, but also significantly improves the detection performance of occluded targets and reduces the missed detection rate caused by occlusion. By improving the feature fusion strategy, the detection accuracy is improved, making the model more robust in complex underwater environments.
[0093] The present embodiment also provides a complex underwater target detection system based on multi-scale features, comprising:
[0094] an image acquisition device configured to acquire an underwater optical image to be detected and to preprocess the underwater optical image using a relative global histogram stretching algorithm to obtain a preprocessed underwater optical image;
[0095] a model construction device configured to construct a multi-scale feature extraction network;
[0096] a result generation device configured to input the preprocessed underwater optical image into the multi-scale feature extraction network for processing to generate a multi-scale target detection result;
[0097] The multi-scale feature extraction network comprises a feature extraction module and a multi-scale feature fusion module, the feature extraction module extracts multi-scale features through a multi-gradient interaction structure, and the multi-scale feature fusion module integrates the multi-scale features using a progressive fusion strategy.
[0098] The above merely describes a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for detecting complex underwater targets based on multi-scale features, characterized in that, The method comprises the following steps: An underwater optical image to be detected is acquired, and a relative global histogram stretching algorithm is used to pre-process the underwater optical image to obtain a pre-processed underwater optical image; A multi-scale feature extraction network is constructed; The pre-processed underwater optical image is input into the multi-scale feature extraction network for processing to generate a multi-scale target detection result. The multi-scale feature extraction network comprises a feature extraction module and a multi-scale feature fusion module, the feature extraction module extracts multi-scale features through a multi-gradient interaction structure, and the multi-scale feature fusion module integrates the multi-scale features using a progressive fusion strategy.
2. The method for detecting complex underwater targets based on multi-scale features according to claim 1, characterized in that, The pre-processing of the underwater optical image using the relative global histogram stretching algorithm comprises the following steps: The RGB channels of the underwater optical image are decomposed into independent channels and subjected to color balance correction; The corrected image channels are subjected to dynamic histogram stretching, and a bilateral filter is used to eliminate noise and retain details; The L component is subjected to linear stretching in the CIE Lab color space, and the a and b components are subjected to S-shaped curve correction, respectively, and finally converted back to the RGB model to obtain the processed underwater optical image.
3. The method for detecting complex underwater targets based on multi-scale features according to claim 2, characterized in that, The RGB channels of the underwater optical image are decomposed into independent channels and subjected to color balance correction, which comprises the following steps: In the formula, R avg , G avg , and B avg respectively represent the normalized average values of the recovered red, green, and blue channels, M and N both represent the spatial resolution of the image, I g and I b respectively represent the gray value of green and the gray value of blue of each pixel point, i is the column where the pixel is located, j is the row where the pixel is located, θ g is the color equalization coefficient of the G channel, and θ b is the color equalization coefficient of the B channel.
4. The method for detecting complex underwater targets based on multi-scale features according to claim 2, characterized in that, The dynamic histogram stretching is performed as follows: where p in and p out represent input and output pixels, respectively, I min , I max , Q min and Q max are adaptive parameters of the images before and after stretching.
5. The method for detecting complex underwater targets based on multi-scale features according to claim 2, characterized in that, The linear stretching of the L component in the CIE Lab color space is performed as follows: In the formula, RD is the stretching result of the pixel value, x is the gray value, and a is the darkest gray value; The S-shaped curve correction is performed as follows: where p χ is the output pixel, I χ is the input pixel, a is the darkest gray, and b is the brightest gray.
6. The method for detecting complex underwater targets based on multi-scale features according to claim 1, characterized in that, The feature extraction module comprises a backbone network and a multi-gradient interaction structure, wherein the backbone network extracts multi-scale features using a convolutional neural network, the multi-gradient interaction structure is composed of three parallel branches, the features are interacted through a disordering operation after channel splicing, and a SimAM attention mechanism is used for semantic feature weighting.
7. The method of claim 6, wherein the method further comprises: The multi-gradient interaction structure comprises: A first branch comprising a 1*1 convolution for extracting first-layer features; A second branch comprising a 1*1 convolution and four 3*3 convolution blocks, and two gradient paths are introduced simultaneously, for extracting second-layer features; A third branch comprising a 1*1 convolution and four 5*5 convolutions, and four gradient paths are introduced simultaneously, for extracting third-layer features.
8. The method of claim 7, wherein the method further comprises: The multi-scale feature fusion module sequentially fuses the first-layer features, the second-layer features and the third-layer features through an adaptive feature pyramid network (AFPN), performs upsampling and downsampling operations to unify the feature map size, and realizes cross-scale adaptive spatial fusion through an ASFF3 unit.
9. The method of claim 8, wherein the method further comprises: In the feature fusion process, an adaptive weight distribution mechanism is used to dynamically adjust the contribution of different scale features according to the feature response intensity.
10. A multi-scale feature based complex underwater target detection system, characterized in that, The method comprises the following steps: An image acquisition device is used to acquire an underwater optical image to be detected, and a relative global histogram stretching algorithm is used to pre-process the underwater optical image to obtain a pre-processed underwater optical image; A model construction device is used to construct a multi-scale feature extraction network; The result generation device is configured to input the preprocessed underwater optical image into the multi-scale feature extraction network for processing to generate a multi-scale target detection result. The multi-scale feature extraction network comprises a feature extraction module and a multi-scale feature fusion module. The feature extraction module extracts multi-scale features through a multi-gradient interaction structure. The multi-scale feature fusion module integrates the multi-scale features using a progressive fusion strategy.