Multi-scale feature fusion polar region unmanned ship target identification method, system and equipment

By adopting a multi-scale feature fusion method in polar unmanned boat target recognition, and using improved E-BiFPN and optimized D-Head detection heads, the missed detection, high error detection rate and semantic information loss of target detection tasks in polar environments are solved, achieving higher positioning accuracy and recognition accuracy.

CN119942330APending Publication Date: 2025-05-06HARBIN ENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510015571.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In polar environments, optical object detection tasks are disturbed by complex environmental factors, resulting in high missed and missed detection rates. The existing lightweight object detection architecture loses semantic information and local detailed features when the network deepens, affecting positioning accuracy and recognition accuracy.

Method used

The multi-scale feature fusion polar unmanned boat target recognition method is adopted. Through the improved E-BiFPN and the optimized D-Head detection head, combined with the Node feature information stacking module and the dynamic detection head, the multi-scale feature fusion capability and robustness of the model are improved.

Benefits of technology

The positioning accuracy and recognition accuracy of polar unmanned boat target recognition are improved, and the adaptability and efficiency of the model to target detection tasks in polar environments are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942330A_ABST
    Figure CN119942330A_ABST
Patent Text Reader

Abstract

The invention discloses a polar region unmanned ship target identification method, system and equipment based on multi-scale feature fusion, belongs to the technical field of target identification, and solves the problems of high missing report rate and weak positioning robustness when an optical sensor performs a target identification task during autonomous navigation. The method comprises the following steps: aiming at a YOLO target detection algorithm, providing a lighter and more efficient E-BiFPN, and fusing more effective scale feature information in a training process by the structure of the E-BiFPN; a lightweight feature information stacked node module is designed. According to the scheme, separable convolution and channel shuffling thoughts are combined; a dynamic detection head is combined, the adaptability of an original model to a water surface target detection task is enhanced, a target detection head and attention are unified, and multiple attention mechanisms, scale consciousness between feature levels, space consciousness between space positions and task consciousness in an output channel are combined in a coordinated and consistent mode. The method is suitable for target identification of the polar region unmanned ship in a severe polar region environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of target recognition, and in particular to polar unmanned boat target recognition. Background Art

[0002] Target detection technology is one of the key technologies for intelligent perception of unmanned boats. Among them, surface optical target detection based on deep learning has been widely studied. Visible light images are rich in information, easy to obtain, and easy to process. At present, many target detection algorithms can realize surface target detection, but there are still the following difficulties under polar conditions:

[0003] (1) The missed detection and false detection rates are high in complex polar environments. When performing missions in polar unmanned boats, they are often in harsh environments. When performing optical target detection tasks, they will be disturbed by environmental factors such as rain, fog, and strong exposure. Existing detection algorithms are difficult to accurately locate and separate targets in complex scenes, so it is necessary to consider improving the positioning accuracy and robustness of the algorithm.

[0004] (2) The commonly used lightweight target detection architectures suitable for polar unmanned boats can effectively improve some speeds, but as the network deepens, a lot of semantic information and local detail features will be lost, resulting in a large loss of accuracy.

[0005] Polar unmanned boats have different requirements for positioning accuracy and recognition accuracy, because unmanned boats are more difficult to control than unmanned vehicles. The polar environment is more open than the common environment, and polar unmanned boats need to operate in harsh environments with floating ice obstructions, overexposure, and fog. In addition, data collection in the polar unmanned boat working environment is more difficult, and the collected data sets are difficult to cover the various complex situations in the surface working environment. Therefore, surface target recognition in the polar environment is more difficult than normal target recognition. Summary of the invention

[0006] The purpose of the present invention is to solve the problems of floating ice, strong background reflection and so on in the polar navigation environment, which lead to high missed reporting rate and weak positioning robustness of optical sensors when performing target recognition tasks during autonomous navigation. A polar unmanned boat target recognition method, system and equipment with multi-scale feature fusion are provided.

[0007] The present invention is implemented by the following technical solutions. On the one hand, the present invention provides a polar unmanned boat target recognition method based on multi-scale feature fusion, the method comprising:

[0008] Step 1: Acquire surface target data;

[0009] Step 2: Input the surface target data in step 1 into the YOLO target detection model based on the improved E-BiFPN and optimized D-Head detection head. The target detection model specifically includes:

[0010] The original feature pyramid is replaced by an improved E-BiFPN, wherein the improved E-BiFPN specifically includes:

[0011] A fusion method based on the Node feature information stacking module and adaptive weights is used as channel splicing. A gradient parameter is defined in the input feature map of each different network layer, and adaptive learning is performed during the training process. After normalization, the input feature maps are multiplied with the respective input feature maps, and finally the feature maps of each network layer are added to obtain the final output. Skip links are introduced to reuse medium-scale feature information that appears more than a preset value in surface tasks. The fusion of shallow feature maps is added in the downsampling part to compensate for the loss of semantic information of small targets in polar scenes.

[0012] Replace the original YOLO detection head with a dynamic detection head;

[0013] Step 3: Visualize and compare the detection results. Visualize the detection results of the water surface data.

[0014] Furthermore, the improved E-BiFPN is an improvement based on BiFPN, and the BiFPN specifically includes:

[0015] Multiple nodes are used to represent features at different levels. Each node has its own weight, which is adjusted dynamically according to the quality and importance of the feature. The weight is calculated using the softmax function for normalization, where the larger the weight value, the greater the contribution. The calculation method is shown in the formula:

[0016]

[0017] Where O represents the output feature map after feature fusion, w represents the adaptively learned weight, I represents the original input feature map, and ε is the minimum value of 0.0001 set to prevent unstable data calculation.

[0018] Furthermore, the Node feature information stacking module infiltrates the information generated by CNN into each information part generated by DSC through channel shuffling; the channel shuffling adopts a uniform mixing strategy, and by uniformly exchanging local feature information on different channels, the information of CNN is fully mixed into the output of DSC.

[0019] Furthermore, the Node feature information stacking module includes a GS Bottleneck module, and the GS Bottleneck module includes two GSConv layers for optimizing convolution calculations.

[0020] Furthermore, the dynamic detection head specifically includes:

[0021] Combined with DCNv4 to optimize the computational offset and dynamic weight in the spatial perception attention of the dynamic detection head, the D-Head is reconstructed to achieve structural optimization and reduce its memory access cost.

[0022] Furthermore, DCNv4 includes: canceling the softmax normalization in spatial aggregation, converting the regulation scalar ranging from [0,1] to unbounded dynamic weights similar to convolution, and merging the linear layers for calculating offsets and dynamic weights into one linear layer.

[0023] Furthermore, DCNv4 specifically includes: optimizing memory access, when using CUDA to execute the previous DCN, for the shape (H, W, C), offset (H, W, G, K 2 ×2) and aggregation weights (H, W, G, K 2 ), a total of H×W×C threads will be created to maximize parallelism, where each thread processes one channel of one output location.

[0024] In a second aspect, the present invention provides a polar unmanned boat target recognition system with multi-scale feature fusion, the system comprising:

[0025] Data acquisition module, used to obtain surface target data;

[0026] The model detection module is used to input the surface target data into the YOLO target detection model based on the improved E-BiFPN and the optimized D-Head detection head. The target detection model specifically includes:

[0027] The original feature pyramid is replaced by an improved E-BiFPN, wherein the improved E-BiFPN specifically includes:

[0028] A fusion method based on the Node feature information stacking module and adaptive weights is used as channel splicing. A gradient parameter is defined in the input feature map of each different network layer, and adaptive learning is performed during the training process. After normalization, the input feature maps are multiplied with the respective input feature maps, and finally the feature maps of each network layer are added to obtain the final output. Skip links are introduced to reuse medium-scale feature information that appears more than a preset value in surface tasks. The fusion of shallow feature maps is added in the downsampling part to compensate for the loss of semantic information of small targets in polar scenes.

[0029] Replace the original YOLO detection head with a dynamic detection head;

[0030] The visualization module is used for visual comparison of detection results and visualization of the detection results of water surface data.

[0031] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the steps of the polar unmanned boat target recognition method with multi-scale feature fusion as described above are executed.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a plurality of computer instructions are stored, and the plurality of computer instructions are used to enable a computer to execute a polar unmanned boat target recognition method with multi-scale feature fusion as described above.

[0033] Beneficial effects of the present invention:

[0034] The present invention provides an optical image detection method for polar unmanned boats based on multi-scale feature fusion, including an E-BiFPN feature fusion pyramid structure for polar unmanned boats, which fully utilizes the advantageous characteristics of a deep neural network model to achieve more accurate perception and construct polar environmental situation.

[0035] Aiming at the problem of insufficient scale feature extraction when the YOLO target detection algorithm faces polar background interference, the present invention proposes a more lightweight and efficient E-BiFPN. This structure integrates more effective scale feature information during the training process, reduces the model calculation amount and improves the speed.

[0036] The present invention designs a node module for lightweight feature information stacking, which combines separable convolution with channel shuffling. Compared with the original network structure, it can stack scale feature information in a simpler way, is more suitable for the deployment and efficient operation of polar unmanned boat target detection algorithms, and effectively improves the environmental perception performance of polar unmanned boats.

[0037] The present invention combines a dynamic head (D-Head) to enhance the adaptability of the original model to the task of surface target detection. The target detection head is unified with attention, and multiple attention mechanisms, scale awareness between feature levels, spatial awareness between spatial positions, and task awareness in the output channel are combined in a coordinated manner.

[0038] The present invention carries out research on optical target perception technology in polar environments, optimizes the feature extraction capability of the model, and takes into account the method deployment requirements on the device; and the present invention combines the characteristics of polar scenes to improve the YOLO network with a more efficient network structure to improve the performance of the target detection algorithm in polar scenes.

[0039] The present invention combines the typical characteristics of the polar unmanned boat working environment and uses multi-scale feature fusion technology to improve the target recognition effect in harsh polar environments, thereby realizing the positioning and high-accuracy recognition of targets by unmanned boats in polar environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0041] Figure 1 It is a schematic diagram of the architecture of the polar unmanned boat target recognition method based on multi-scale feature fusion of the present invention;

[0042] Figure 2 It is a structural comparison diagram of BiFPN of the present invention, FPN, PANet and NAS-FPN;

[0043] Figure 3 This is a schematic diagram of replacing FPN with E-BiFPN in the present invention;

[0044] Figure 4 This is a schematic diagram of the principle of channel shuffling;

[0045] Figure 5 Output visualizations of feature maps for different convolutions;

[0046] Figure 6 This is the structure diagram of GSConv and Node node modules;

[0047] Figure 7 This is the D-Head network structure diagram;

[0048] Figure 8 Schematic diagram for optimizing memory access;

[0049] Fig. 9 Comparison chart of detection effects in different scenarios;

[0050] Fig.10 It is a heat map visualization comparison chart;

[0051] Fig.11 This is a visualization comparison of the effective receptive field. DETAILED DESCRIPTION

[0052] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0053] Embodiment 1: A polar unmanned boat target recognition method using multi-scale feature fusion, the method comprising:

[0054] Step 1: Acquire surface target data;

[0055] Step 2: Input the surface target data in step 1 into the YOLO target detection model based on the improved E-BiFPN and optimized D-Head detection head. The target detection model specifically includes:

[0056] The original feature pyramid is replaced by an improved E-BiFPN, wherein the improved E-BiFPN specifically includes:

[0057] A fusion method based on the Node feature information stacking module and adaptive weights is used as channel splicing. A gradient parameter is defined in the input feature map of each different network layer, and adaptive learning is performed during the training process. After normalization, the input feature maps are multiplied with the respective input feature maps, and finally the feature maps of each network layer are added to obtain the final output. Skip links are introduced to reuse medium-scale feature information that appears more than a preset value in surface tasks. The fusion of shallow feature maps is added in the downsampling part to compensate for the loss of semantic information of small targets in polar scenes.

[0058] It should be noted that the fusion method combining feature information stacking and adaptive weight is an optimized adaptive weight fusion method, which combines node and adaptive weight.

[0059] Replace the original YOLO detection head with a dynamic detection head;

[0060] Step 3: Visualize and compare the detection results. Visualize the detection results of the water surface data.

[0061] This implementation method studies polar target detection based on the YOLO target detection algorithm. In view of the fact that target features in polar scenes are interfered by background, an improved E-BiFPN is proposed to replace the original feature pyramid network (FPN), and combined with a dynamic head (D-Head), the adaptability of the original model to the task of surface target detection is enhanced, thereby enhancing the model's ability to fuse multi-scale features.

[0062] Small target detection performance is an important indicator for evaluating the long-distance perception capability of unmanned platforms. For example, tasks such as surface surveillance patrols and dynamic tracking rely on this performance. Polar target scenes are complex and vast. Small targets such as small boats and ice floes occupy limited pixel information in the image, so they are more difficult to detect in harsh environments such as rain, fog and strong reflections. To solve these problems, this implementation further draws on the idea of ​​BiFPN to reconstruct the feature pyramid, and proposes an efficient bidirectional weighted feature pyramid that is lighter and can effectively extract more scale features, referred to as E-BiFPN. Among them, the idea of ​​BiFPN can be easily applied to the pyramid structure of target detection, and has a certain enhanced feature extraction capability.

[0063] Implementation method 2: This implementation method further limits the polar unmanned boat target recognition method based on multi-scale feature fusion as described above. In this implementation method, the improved E-BiFPN is further limited, specifically including:

[0064] The improved E-BiFPN is an improvement based on BiFPN, and the BiFPN specifically includes:

[0065] Multiple nodes are used to represent features at different levels. Each node has its own weight, which is adjusted dynamically according to the quality and importance of the feature. The weight is calculated using the softmax function for normalization, where the larger the weight value, the greater the contribution. The calculation method is shown in the formula:

[0066]

[0067] Where O represents the output feature map after feature fusion, w represents the adaptively learned weight, I represents the original input feature map, and ε is the minimum value of 0.0001 set to prevent unstable data calculation.

[0068] The BiFPN of this embodiment is a deep learning structure for target detection. BiFPN aims to solve the problems of information loss and insufficient information transfer in the traditional unidirectional feature pyramid network (FPN). BiFPN introduces a bidirectional information flow, which enables the network to better fuse features at different levels. The traditional FPN structure is unidirectional and can only transfer feature information from low layers to high layers. BiFPN introduces a bidirectional information flow, which can not only transfer information from low layers to high layers, but also from high layers to low layers. This bidirectional information flow helps to better utilize the features of different layers in the network, and can more comprehensively capture the multi-scale information of the target. Secondly, a dynamic feature fusion mechanism is introduced in BiFPN, which can adaptively fuse features at different levels.

[0069] This embodiment adopts the BiFPN because BiFPN introduces bidirectional information flow and dynamic feature fusion, but its design is still relatively lightweight with low computational and parameter complexity. This enables BiFPN to achieve excellent performance in target detection tasks and is suitable for resource-constrained scenarios such as mobile terminals and embedded devices. By introducing innovations such as bidirectional information flow and dynamic feature fusion, BiFPN effectively solves the problems of information loss and insufficient information transmission in traditional unidirectional feature pyramid networks, thereby improving the performance and efficiency of target detection networks.

[0070] Implementation method three: This implementation method is a further limitation on the polar unmanned boat target recognition method based on multi-scale feature fusion as described above. In this implementation method, the Node feature information stacking module is further limited, specifically including:

[0071] The Node feature information stacking module infiltrates the information generated by CNN into each information part generated by DSC through channel shuffling; the channel shuffling adopts a uniform mixing strategy, and by uniformly exchanging local feature information on different channels, the information of CNN is fully mixed into the output of DSC.

[0072] This embodiment designs a Node feature information stacking module for use in the E-BiFPN structure, which adopts a construction method based on Ghost-Shuffle Convolutional (GSConv). For intelligent unmanned devices, such as unmanned boats, speed and accuracy are both critical. Traditional lightweight architectures such as Xception, MobileNet, and GhostNet have increased speed. However, when applied to the field of autonomous driving, these models generally exhibit lower accuracy. These structures are designed to accelerate the model prediction speed, but due to the sparseness of their feature sampling methods, the simplification of inference operator steps, and other operations, each time the spatial size of the feature map is transformed and channel information is fused, a part of the semantic information will be lost more or less. However, GSConv retains these connections as much as possible with a lower time complexity. GSConv draws on the idea of ​​channel shuffling in ShuffleNet, and by combining channels of different groups together, it makes up for the shortcomings of point-by-point group convolution, allowing information between groups to flow freely.

[0073] GSConv organically combines the ideas of CNN, Deepwise Seperable Convolution (DSC) and channel shuffling, and uses channel shuffling to infiltrate the information generated by CNN into each part of the information generated by DSC. Channel shuffling adopts a uniform mixing strategy, which uniformly exchanges the local feature information on different channels, so that the information of CNN is fully mixed into the output of DSC without introducing redundant information.

[0074] Implementation method 4: This implementation method is a further limitation on the polar unmanned boat target recognition method based on multi-scale feature fusion as described above. In this implementation method, the Node feature information stacking module is further limited, specifically including:

[0075] The Node feature information stacking module includes a GS Bottleneck module, which includes two GSConv layers for optimizing convolution calculations.

[0076] In the structure of this implementation, Node is the top-level module, and GS Bottleneck, as the core part, reduces computational complexity through the GSConv layer and improves the computational efficiency of the model. This design helps optimize the performance of the network by introducing bottleneck structures and grouped convolutions, especially when processing large-scale data sets, which can effectively control the consumption of computing resources while maintaining high expressiveness and computational efficiency.

[0077] In general, GSConv builds a simpler gradient information stacking method compared to the original YOLO structure, and is especially lighter than the multi-branch stacking structure in the original model, but reduces the feature reuse rate. However, for edge deployment hardware friendliness such as polar unmanned boats, a simpler and more efficient structure is easier to use.

[0078] Embodiment 5: This embodiment further limits the polar unmanned boat target recognition method based on multi-scale feature fusion as described above. In this embodiment, the dynamic detection head is further limited, specifically including:

[0079] The dynamic detection head specifically comprises:

[0080] Combined with DCNv4 to optimize the computational offset and dynamic weight in the spatial perception attention of the dynamic detection head, the D-Head is reconstructed to achieve structural optimization and reduce its memory access cost.

[0081] This implementation reconstructs the D-Head, cancels the softmax in DCNv4, converts the adjustment scalar in the range of [0,1] into an unbounded dynamic weight similar to convolution, and merges the linear layers for calculating offsets and dynamic weights into one linear layer. This change further expands the dynamic characteristics of DCN, making the convergence speed of DCNv4 significantly faster than DCNv3 and other common operators, including convolution and attention. In addition, DCNv4 also significantly improves the speed of DCN by saving additional memory instructions and makes the speed advantage of sparse operators a reality.

[0082] Implementation method 6: This implementation method further limits the polar unmanned boat target recognition method based on multi-scale feature fusion as described above. In this implementation method, DCNv4 is further limited, specifically including:

[0083] DCNv4 includes: canceling the softmax normalization in spatial aggregation, converting the adjustment scalar ranging from [0,1] to unbounded dynamic weights similar to convolution, and merging the linear layers for calculating offsets and dynamic weights into one linear layer.

[0084] This implementation eliminates the softmax normalization in spatial aggregation to enhance its dynamic characteristics and expressiveness. This improvement further expands the dynamic characteristics of DCN, making the convergence speed of DCNv4 significantly faster than DCNv3 and other common operators, including convolution and attention.

[0085] Implementation method seven: This implementation method further limits the polar unmanned boat target recognition method based on multi-scale feature fusion as described above. In this implementation method, DCNv4 is further limited, specifically including:

[0086] DCNv4 specifically includes: optimizing memory access, using CUDA to execute the previous DCN for data with shape (H, W, C), offset (H, W, G, K 2 ×2) and aggregation weights (H, W, G, K 2 ), a total of H×W×C threads will be created to maximize parallelism, where each thread processes one channel of one output location.

[0087] DCNv4 in this implementation addresses the limitations of previous DCN versions and makes two key improvements: first, the softmax normalization in spatial aggregation is cancelled to enhance its dynamic characteristics and expressiveness; then, memory access is optimized to minimize redundant operations to increase speed.

[0088] Embodiment 8: This embodiment is an example of a polar unmanned boat target recognition method using multi-scale feature fusion as described above, and specifically includes:

[0089] Based on the YOLO target detection algorithm, the polar target detection problem is studied. In view of the fact that target features in polar scenes are interfered by background, an improved E-BiFPN is proposed to replace the original feature pyramid (FeaturePyramidNetwork, FPN), combined with the dynamic detection head (Dynamic Head, D-Head), to enhance the adaptability of the original model to the task of surface target detection, thereby enhancing the model's ability to fuse multi-scale features. The architecture of the method of the present invention is as follows Figure 1 shown.

[0090] (1) Combined with bidirectional weighted feature pyramid:

[0091] Combining the idea of ​​BiFPN (Bidirectional Feature Pyramid Network) to further improve the FPN results of YOLOv7. Its structure is as follows Figure 2 As shown in d. Figure 2 As shown, the structure of the BiFPN of the present invention is different from that of the prior art FPN, PANet and NAS-FPN.

[0092] BiFPN aims to solve the problems of information loss and insufficient information transmission in the traditional unidirectional feature pyramid network (FPN). Specifically, it uses multiple nodes to represent features at different levels. Each node has its own weight, and the weight is dynamically adjusted according to the quality and importance of the feature, thereby achieving more flexible and efficient feature fusion. The weight calculation is normalized using the softmax function, where the larger the weight value, the greater the contribution, and the obtained feature is more valued by the Neck network. The calculation method is shown in the formula:

[0093]

[0094] Where O represents the output feature map after feature fusion, w represents the adaptively learned weight, I represents the original input feature map, and ε is the minimum value of 0.0001 set to prevent unstable data calculation.

[0095] (2) Efficient bidirectional weighted feature pyramid design:

[0096] Small target detection performance is an important indicator for evaluating the long-distance perception capability of unmanned platforms. Tasks such as surface surveillance patrols and dynamic tracking rely on this performance. Polar target scenes are complex and vast. Small targets such as small boats and ice floes occupy limited pixel information in the image, and are therefore more difficult to detect in harsh environments such as rain, fog, and strong reflections. To solve these problems, the present invention further draws on the idea of ​​BiFPN to reconstruct the feature pyramid, and proposes an efficient bidirectional weighted feature pyramid that is lighter and can effectively extract more scale features, referred to as E-BiFPN. Figure 3 shown.

[0097] The idea of ​​BiFPN can be easily applied to the pyramid structure of target detection, and has a certain ability to enhance feature extraction. E-BiFPN uses an optimized Weight (adaptive weight) fusion method to replace the original channel splicing (ConCat), defines a gradient parameter in the input feature map of each different network layer, and performs adaptive learning during the training process. After normalization, it is multiplied with the respective input feature maps, and finally the feature maps of each network layer are added to obtain the final output. Secondly, skip links are introduced to reuse the medium-scale feature information that appears more frequently in surface tasks. Further, the fusion of shallow feature maps is added in the downsampling part to compensate for the loss of semantic information of small targets in polar scenes in this process.

[0098] (3) Lightweight information stacking module design:

[0099] A Node feature information stacking module is designed for use in the E-BiFPN structure. It is constructed based on the Ghost-Shuffle Convolutional (GSConv) method. GSConv combines channels from different groups together to make up for the shortcomings of point-by-point group convolution, allowing information between groups to flow freely. Its principle is as follows Figure 4 shown.

[0100] GSConv organically combines the ideas of CNN, Depthwise Seperable Convolution (DSC) and channel shuffling, and infiltrates the information generated by CNN into each information part generated by DSC through channel shuffling. Channel shuffling adopts a uniform mixing strategy, which makes the information of CNN fully mixed into the output of DSC by uniformly exchanging the local feature information on different channels without introducing redundant information.

[0101] Figure 5The feature map visualization results of CNN, DSC and GSConv are shown. It can be clearly observed from the figure that the feature map of GSConv has a higher similarity than that of CNN, and its similarity is also significantly improved compared with DSC. This shows that as the depth of the network model increases, GSConv retains the feature information more effectively.

[0102] The Node module contains a GS Bottleneck module. The GS Bottleneck module further contains two GSConv layers for optimizing convolution calculations. The GSConv and Node module structures are shown in the figure below. Figure 6 As shown. In general, GSConv builds a simpler gradient information stacking method compared to the original YOLO structure, especially lighter than the multi-branch stacking structure in the original model, but reduces the feature reuse rate. However, for edge deployment hardware friendliness such as polar unmanned boats, a simpler and more efficient structure is easier to use.

[0103] (4) Dynamic detection head optimization:

[0104] The original YOLO detection head is replaced by a dynamic head, and the number of channels of the predicted output is uniformly adjusted to 128. The network structure of the dynamic detection head D-Head is as follows Figure 7 shown.

[0105] Combined with Deformable Convolution v4, namely DCNv4, to optimize the computational offset and dynamic weight in the spatial perception attention of the D-Head, the D-Head is further reconstructed to achieve structural optimization and reduce its memory access cost.

[0106] DCNv4 addresses the limitations of previous DCN versions and makes two key improvements: first, the softmax normalization in spatial aggregation is cancelled to enhance its dynamic characteristics and expressiveness; then memory access is optimized to minimize redundant operations to improve speed. Taking DCNv3 as an example, it can be seen as a combination of convolution and attention. On the one hand, it processes input data in a sliding window manner, which follows convolution and its inductive bias; on the other hand, the sampling offset Δp and the spatial aggregation weight m are dynamically predicted based on the input features, showing its dynamic characteristics, making it more like an attention mechanism. Comparing different operators, a major difference between convolution and DCNv3 is that DCNv3 uses a softmax function to normalize the spatial aggregation weight m, following the convention of self-attention. In contrast, CNN does not use it on weights, but still works well. The reason why attention needs to use softmax is that Q, K, V∈R N×d The self-attention of is defined by the following formula:

[0107]

[0108] Where N is the number of points in the same attention window (can be global or local), d is the hidden dimension, and Q, K, V are the query, key, and value matrices computed from the input.

[0109] In the above formula, if there is no softmax operation, KV∈R d×d It can be calculated first that for all queries in the same attention window, it will degenerate into a linear projection, resulting in poor performance. In fact, using softmax to normalize the convolution weights in a fixed range of [0,1] will greatly limit the expressiveness of the operator and slow down the learning speed. In DCNv4, softmax is cancelled, and the adjustment scalar in the range of [0,1] is converted to an unbounded dynamic weight similar to convolution, and the linear layers for calculating offsets and dynamic weights are merged into one linear layer. This change further expands the dynamic characteristics of DCN, making DCNv4 converge significantly faster than DCNv3 and other common operators, including convolution and attention.

[0110] In addition, DCNv4 significantly improves the speed of DCN by saving additional memory instructions and turns the speed advantage of sparse operators into reality, such as Figure 8 shown.

[0111] When using CUDA to execute the previous DCN, for the shape (H, W, C), offset (H, W, G, K 2 ×2) and aggregation weights (H, W, G, K 2 ), a total of H×W×C threads will be created to maximize parallelism, where each thread processes one channel of one output position. It is worth noting that the D=C / G channels in each group share the same sampling offset and aggregation weight value for each output position. Using multiple threads to process these D channels on the same output position is wasteful because different threads will read the same sampling offset and aggregation weight value from GPU memory multiple times, which is critical for memory-constrained operations. Using one thread to process multiple channels in the same group on each output position can eliminate these redundant memory read requests, thereby greatly reducing memory usage. Since the sampling positions are the same, if each thread processes D 1 channels, then the memory access cost of reading the offset and aggregation weights and the computational cost of calculating the bilinear interpolation coefficients can be reduced by D 1 times.

[0112] (5) Visual verification:

[0113] (5.1) Visual comparison of test results:

[0114] The model is visualized to demonstrate its effectiveness in multiple aspects. First, foggy days, small targets, target occlusion, weak targets and other scenes are selected from a variety of complex water scenes in the test set for comparison of detection effects. Fig. 9 As shown in the figure, the Ours model is on the left and the Base model is on the right. The green dashed box represents the detection effect of the improved Ours model, and the red dashed box represents the detection effect of the original model. It can be seen that the integrated solution has better capabilities than the baseline model in focusing on small target features, reusing multi-scale features, and suppressing information loss when long-distance features.

[0115] (5.2) Heat map visualization comparison

[0116] Fig.10 The heat map output by D-Head and the original model decoupling head is compared, which is drawn by the gradient-weighted class activation map (Grad-CAM). The class activation map generated by Grad-CAM can intuitively show the image area that the neural network pays attention to when classifying, thereby helping to understand the decision-making process of the network and providing valuable information for model debugging and optimization.

[0117] From the figure, we can see that the present invention pays more attention to the details of the target, that is, the focused part (red part) is more accurate and covers a wider area. In addition, it also combines more environmental information to focus on small targets that are easily overlooked. These images vividly illustrate that the attention awareness in D-Head enables the model to capture more scale, space, and task information on the deep feature map.

[0118] (5.3) Visualization comparison of effective receptive fields

[0119] like Fig.11 As shown in Figure 1, the effective receptive field in the network is visualized, where the dark green part represents the effective receptive field. Usually, it is necessary to use experimental methods, such as adding small perturbations to specific areas of the input image and then observing the changes in the output to make an approximate estimate.

[0120] Fig.11a is the FPN output of YOLOv8n, and b is the E-BiFPN output of Ours model. Usually, the effective receptive field focuses on the area in the input that has a greater impact on the output. Its size is related to the size k of the convolution kernel and the model depth L, which is proportional to k and inversely proportional to L. The improved model of the present invention has more layers than the original model. However, as the network deepens, the Ours model has a wider effective receptive field before it is input into the D-Head detection head. This shows that the improved E-BiFPN structure of the present invention can provide sufficient receptive layers and the ability to aggregate spatial information. The representation ability of the model is also better in deep networks, providing more nonlinear and cross-channel information exchange.

Claims

1. A polar unmanned boat target recognition method based on multi-scale feature fusion, characterized in that: The method comprises: Step 1: Acquire surface target data; Step 2: Input the surface target data in step 1 into the YOLO target detection model based on the improved E-BiFPN and optimized D-Head detection head. The target detection model specifically includes: The original feature pyramid is replaced by an improved E-BiFPN, wherein the improved E-BiFPN specifically includes: A fusion method based on the Node feature information stacking module and adaptive weights is used as channel splicing. A gradient parameter is defined in the input feature map of each different network layer, and adaptive learning is performed during the training process. After normalization, the input feature maps are multiplied with the respective input feature maps, and finally the feature maps of each network layer are added to obtain the final output. Skip links are introduced to reuse medium-scale feature information that appears more than a preset value in surface tasks. The fusion of shallow feature maps is added in the downsampling part to compensate for the loss of semantic information of small targets in polar scenes. Replace the original YOLO detection head with a dynamic detection head; Step 3: Visualize and compare the detection results. Visualize the detection results of the water surface data.

2. The polar unmanned boat target recognition method based on multi-scale feature fusion according to claim 1 is characterized in that: The improved E-BiFPN is an improvement based on BiFPN, and the BiFPN specifically includes: Multiple nodes are used to represent features at different levels. Each node has its own weight, which is adjusted dynamically according to the quality and importance of the feature. The weight is calculated using the softmax function for normalization, where the larger the weight value, the greater the contribution. The calculation method is shown in the formula: Where O represents the output feature map after feature fusion, w represents the adaptively learned weight, I represents the original input feature map, and ε is the minimum value of 0.0001 set to prevent unstable data calculation.

3. The polar unmanned boat target recognition method based on multi-scale feature fusion according to claim 1 is characterized in that: The Node feature information stacking module infiltrates the information generated by CNN into each information part generated by DSC through channel shuffling; the channel shuffling adopts a uniform mixing strategy, and by uniformly exchanging local feature information on different channels, the information of CNN is fully mixed into the output of DSC.

4. The polar unmanned boat target recognition method based on multi-scale feature fusion according to claim 3 is characterized in that: The Node feature information stacking module includes a GS Bottleneck module, which includes two GSConv layers for optimizing convolution calculations.

5. The polar unmanned boat target recognition method based on multi-scale feature fusion according to claim 1 is characterized in that: The dynamic detection head specifically comprises: Combined with DCNv4 to optimize the computational offset and dynamic weight in the spatial perception attention of the dynamic detection head, the D-Head is reconstructed to achieve structural optimization and reduce its memory access cost.

6. The method for identifying polar unmanned boat targets by multi-scale feature fusion according to claim 5, characterized in that: DCNv4 includes: canceling the softmax normalization in spatial aggregation, converting the adjustment scalar ranging from [0,1] to unbounded dynamic weights similar to convolution, and merging the linear layers for calculating offsets and dynamic weights into one linear layer.

7. The polar unmanned boat target recognition method based on multi-scale feature fusion according to claim 6 is characterized in that: DCNv4 specifically includes: optimizing memory access, using CUDA to execute the previous DCN for data with shape (H, W, C), offset (H, W, G, K 2 ×2) and aggregation weights (H, W, G, K 2 ), a total of H×W×C threads will be created to maximize parallelism, where each thread processes one channel of one output location.

8. A multi-scale feature fusion polar unmanned boat target recognition system, characterized in that: The system comprises: Data acquisition module, used to obtain surface target data; The model detection module is used to input the surface target data into the YOLO target detection model based on the improved E-BiFPN and the optimized D-Head detection head. The target detection model specifically includes: The original feature pyramid is replaced by an improved E-BiFPN, wherein the improved E-BiFPN specifically includes: A fusion method based on the Node feature information stacking module and adaptive weights is used as channel splicing. A gradient parameter is defined in the input feature map of each different network layer, and adaptive learning is performed during the training process. After normalization, the input feature maps are multiplied with the respective input feature maps, and finally the feature maps of each network layer are added to obtain the final output. Skip links are introduced to reuse medium-scale feature information that appears more than a preset value in surface tasks. The fusion of shallow feature maps is added in the downsampling part to compensate for the loss of semantic information of small targets in polar scenes. Replace the original YOLO detection head with a dynamic detection head; The visualization module is used for visual comparison of detection results and visualization of the detection results of water surface data.

9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor runs the computer program stored in the memory, the steps of the method according to any one of claims 1 to 7 are performed.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of computer instructions, and the plurality of computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Target detection method, device and system for night infrared image

    CN120656126A