Rotary small target real-time detection method, system and device based on multi-scale feature fusion and medium
By using technologies such as multi-scale feature fusion and deformable convolution in infrared images, the problem of insufficient detection accuracy and real-time performance of rotating small targets in complex battlefields is solved, and the accurate identification and efficient detection of infrared rotating small targets is achieved.
Patent Information
- Application Number
- CN202510212479.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has poor accuracy and real-time capability to identify small rotating targets in complex battlefield situations, especially in infrared images, where the target size is small, shape is diverse, and background noise is strong, resulting in increased detection difficulty.
The real-time detection method of rotating small targets based on multi-scale feature fusion is adopted to improve the feature extraction capability of the backbone network through deformable convolution and hollow convolution, and the feature expression of infrared rotating small targets is enhanced through multi-scale feature fusion combined with features of different scales.
In complex scenarios, the accurate identification of infrared rotation small targets is achieved, which improves the accuracy and robustness of detection, reduces the calculation amount and maintains the detection accuracy.
Smart Images

Figure CN120047673A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target recognition, and particularly relates to a real-time detection method, system, device and medium for rotating small targets based on multi-scale feature fusion. Background Art
[0002] Rotating target detection is an important research direction in the field of computer vision, focusing on accurately identifying and detecting rotating targets in images or videos and precisely calculating their rotation angles. This technology is widely used in fields such as military operations, autonomous driving, and drone monitoring. The rotating target detection model receives an image or video containing the target object as input and outputs the bounding box and class label of the target, paying particular attention to the direction and pose of the object to achieve more accurate positioning and recognition. Rotating small target detection, on the other hand, focuses on smaller objects, and the size of small targets may be only a few to dozens of pixels, which makes the detection and recognition of target objects face greater detection difficulties.
[0003] During the process of rotating small target recognition, the input image or video may have a complex background, which may include a cluttered object distribution, resulting in the target being occluded or confused by other objects; shadows and reflections caused by changes in lighting, affecting the visibility of the target; and the similarity in color and texture between the target and the background, making recognition more difficult. In addition, the blur caused by the rapid movement of the target object in real-time detection scenarios and the perspective distortion under different viewpoints will also significantly reduce the recognition accuracy. These factors together increase the challenges of the detection algorithm, requiring the adoption of more advanced rotating small target recognition technologies to improve the robustness and accuracy of recognition.
[0004] For a real-time rotating small target recognition system, real-time performance is also a key indicator. In the fields of computer vision and artificial intelligence, more and more application scenarios require the immediate recognition of fast-moving rotating small targets. Therefore, studying how to improve the real-time performance of the rotating small target detection algorithm while ensuring accuracy has also become one of the important goals of this task.
[0005] Specifically, in future military operations, the real-time detection technology of rotating small targets in complex scenarios has important application prospects in the field of intelligent combat analysis and processing. In future unmanned operations and intelligent cooperative operations, the real-time detection system of rotating small targets can effectively analyze the captured high-quality video image data to identify different types of targets and accurately locate their positions. Ultimately, this will establish a situation awareness system of multi-dimensional information, realize the digitization and networking of battlefield information, comprehensively and accurately grasp the battlefield dynamics, and further support unmanned intelligent weapons such as unmanned aerial vehicles for intelligence reconnaissance, surveillance tracking and precision strikes. By using the automatic real-time detection system of rotating small targets, a large amount of manpower and time costs can be saved, the combat information of our side and the enemy can be grasped in time, and the initiative in the war can be ensured. In addition, the system also has strong anti-interference ability, can timely control the combat information, realize the knowability, controllability and visualization of the global war information, ensure the full exchange and perception of information in the entire battlefield environment, and thus ensure the smooth progress of combat operations.
[0006] With the rapid development and large-scale practical application of unmanned aerial vehicle remote sensing infrared image technology, infrared imaging equipment reconnaissance has become an important means to perceive the battlefield information of the enemy. Infrared imaging technology can accurately capture heat source signals under night and low visibility conditions and identify moving targets of the enemy. Combining with a real-time target detection system, through the dynamic analysis of infrared images, different types of targets can be quickly located and classified, providing accurate data required for tactical decisions. In actual combat, the battlefield environment often includes relatively complex battlefield scenarios, including natural environments and battlefield environments such as cloud clusters, smoke, sea surfaces and complex ground objects, which will cause certain interference to the identification of combat targets such as enemy tanks, armored vehicles, fighter jets, and missiles. At the same time, due to the long imaging distance, real targets are usually small, the signal-to-noise ratio is low, there are few target features that can be extracted, and the targets may also be affected by the background, such as cloud occlusion and ground object interference, resulting in difficult-to-distinguish textures and shapes. Existing domestic and foreign detection algorithms have relatively poor recognition accuracy and real-time performance for such complex battlefield situations. Summary of the Invention
[0007] To overcome the technical problems of relatively poor recognition accuracy and real-time performance for complex battlefield situations in the prior art, the purpose of the present invention is to provide a real-time detection method, system, device and medium for rotating small targets based on multi-scale feature fusion.
[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0009] A real-time detection method for rotating small targets based on multi-scale feature fusion includes the following steps:
[0010] Obtain the infrared image to be detected;
[0011] Extract features from the infrared image to be detected to obtain a feature map;
[0012] Perform multi-scale feature fusion on the feature map to obtain fused features;
[0013] Integrate and classify the fused features to generate the object detection result.
[0014] Furthermore, extracting features from the infrared image to be detected to obtain a feature map includes the following steps:
[0015] Extract features from the infrared image to be detected to obtain feature maps {C_1, C_2, C_3, C_4} of different scales, where C_1 is the first feature map, C_2 is the second feature map, C_3 is the third feature map, and C_4 is the fourth feature map. The extraction process includes the processing of multiple layers of convolution and the CSP layer; the third feature map C_3 performs deformable convolution and dilated convolution to obtain the fifth feature map C_5.
[0016] Furthermore, the length and width of the first feature map C_1 are 1 / 4 of the infrared image to be detected, and the number of channels is 24; the spatial resolutions of the second feature map C_2, the third feature map C_3, and the fourth feature map C_4 are halved in turn, and the number of feature channels doubles layer by layer.
[0017] Furthermore, the third feature map C_3 performs deformable convolution and dilated convolution to obtain the fifth feature map C_5, including the following steps:
[0018] The feature map of the third feature map C_3 after deformable convolution is divided into three paths, each passing through dilated convolution, and then the three paths of features are fused through concatenation operation. The fused feature map extracts and optimizes features through depthwise separable convolution. The third feature map C_3 and the output after extracting and optimizing features through depthwise separable convolution are connected by residual connection to obtain the fifth feature map C_5.
[0019] Furthermore, performing multi-scale feature fusion on the feature map to obtain fused features includes the following steps:
[0020] The second feature map C_2 generates the ninth feature map F_2 through concatenation operation with the upsampling layer, performs local enhancement on the ninth feature map F_2, and then generates the first multi-scale feature map P_2 through convolution module processing;
[0021] The third feature map C_3 forms the sixth feature map T_3 through concatenation and upsampling layer; the sixth feature map T_3 refines features through the LAC layer and convolution module to generate the tenth feature map F_3; the tenth feature map F_3 generates the second multi-scale feature map P_3 through concatenation, CSP layer, and convolution module processing;
[0022] The fourth feature map C_4 passes through the convolutional module to obtain the eleventh feature map F_4. The features after concatenating the ninth feature map F_2 and the tenth feature map F_3 are concatenated with the features of the CSP layer and the convolutional module, and then concatenated with the eleventh feature map F_4 and passed through the CSP layer to obtain the third multi-scale feature map P_4;
[0023] The fifth feature map C_5 is adjusted by the channel convolutional module to generate the fourth multi-scale feature map P_5.
[0024] Further, in the LAC layer, first perform a convolutional operation on the sixth feature map T_3 or the ninth feature map F_2, combine with the CSP layer for feature extraction and channel decomposition, and then integrate the multi-level information through concatenation and convolutional operations to obtain the integrated feature map F; the integrated feature map F passes through the LAC layer to obtain the seventh feature map U; the seventh feature map U optimizes the information flow through the gating mechanism to obtain the eighth feature map Y; among them, the gating mechanism extracts the global statistical information of the features through global average pooling, and is processed through convolution and activation functions to generate the weights for feature selection; the weights for feature selection are applied to the seventh feature map U through element-wise multiplication, and finally the original feature information is retained through residual connection to obtain the eighth feature map Y.
[0025] Further, integrate and classify the fused features to generate the object detection results, including the following steps:
[0026] The first multi-scale feature map P_2, the second multi-scale feature map P_3, the third multi-scale feature map P_4, and the fourth multi-scale feature map P_5 pass through the shared convolutional layer and are then passed to the classification branch and the regression branch;
[0027] In the classification branch, each feature map is processed by batch normalization and then fed into the classification head for class prediction. The class prediction results pass through the Softmax function or the Sigmoid function to generate the class distribution of the object;
[0028] In the regression branch, each feature map is processed by batch normalization and then fed into the regression head to calculate the bounding box position of the object.
[0029] A real-time detection system for rotating small objects based on multi-scale feature fusion, including:
[0030] An image acquisition module for acquiring the infrared image to be detected;
[0031] A feature extraction module for extracting features from the infrared image to be detected to obtain feature maps;
[0032] A multi-scale feature fusion module for performing multi-scale feature fusion on the feature maps to obtain the fused features;
[0033] An integration and classification module is used to integrate and classify the fused features to generate object detection results.
[0034] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the real-time detection method for rotating small targets based on multi-scale feature fusion is implemented.
[0035] A computer-readable storage medium stores a computer program. The computer program, when executed by a processor, implements the real-time detection method for rotating small targets based on multi-scale feature fusion.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] Aiming at the characteristics of complex scenarios, the present invention improves the feature extraction ability of the backbone network by adopting deformable convolution and dilated convolution. At the same time, through multi-scale feature fusion, features of different scales are effectively combined, enhancing the feature expression of infrared rotating small targets and achieving accurate recognition of infrared rotating small targets in complex scenarios. The present invention not only improves the technical level of infrared rotating small target detection, but also provides strong technical support for the field of intelligent combat analysis and processing. Through the real-time detection system, high-quality video image data can be effectively analyzed, different types of targets can be identified and their positions accurately located, a multi-dimensional information situation awareness system can be established, the battlefield information can be digitalized and networked, and the battlefield dynamics can be comprehensively and accurately grasped. In the military field, infrared small target detection technology can conduct reconnaissance and surveillance on distant targets, such as unmanned aerial vehicles, missiles, etc.; in the security field, infrared small target detection technology can timely detect abnormal targets and improve public security. This will greatly enhance the intelligence reconnaissance, surveillance tracking, and precise strike capabilities of unmanned intelligent weapons such as unmanned aerial vehicles, save a large amount of manpower and time costs, and ensure the initiative in war. In addition, with the rapid development of unmanned aerial vehicle remote sensing infrared image technology, infrared imaging device reconnaissance has become an important means to perceive enemy battlefield information. The research results of the present invention will effectively enhance the combat effectiveness of the military, play a unique and irreplaceable role, and have important significance for improving the national security defense ability and coping with complex and changeable security environments.
[0038] Furthermore, the present invention adopts depthwise separable convolution and channel compression technology to reduce the computational amount while maintaining the detection accuracy.
[0039] Furthermore, when performing multi-scale fusion, multiple upsamplings and CSP layers are adopted to enrich the expression ability of the feature map. Description of the Drawings
[0040] Figure 1It is the structural diagram of the improved RTMDet network;
[0041] Figure 2 It is the structural diagram of the deformable dilated convolution module;
[0042] Figure 3 It is the structural diagram of the LAC layer, gating, and convolution module;
[0043] Figure 4 It is the structural diagram of the LAC layer;
[0044] Figure 5 It is the flowchart of strong data augmentation;
[0045] Figure 6 It is the diagram of the caching mechanism;
[0046] Figure 7 It is the diagram of depthwise separable convolution;
[0047] Figure 8 It is the structural diagram of the DC convolution module;
[0048] Figure 9 It is the overall flowchart of the real-time detection method for rotated small targets based on multi-scale feature fusion;
[0049] Figure 10 It is the flowchart of the real-time detection method for rotated small targets based on multi-scale feature fusion;
[0050] Figure 11 It is the schematic diagram of the real-time detection system for rotated small targets based on multi-scale feature fusion;
[0051] Figure 12 It is the schematic diagram of the actual hardware deployment of the real-time detection system for rotated small targets based on multi-scale feature fusion. Detailed implementation manners
[0052] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. The preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0053] Data-driven deep learning methods can automatically extract robust feature representations from raw data, outperforming traditional extraction methods. In addition, deep learning-based object detection methods can reduce the engineering burden of traditional feature modeling. Deep learning-based object detection frameworks are widely divided into two types: two-stage detection frameworks and one-stage detection frameworks. Two-stage detection methods use the selective search algorithm to extract candidate regions in the first stage. Then, CNN is used to extract features from the candidate regions and a classifier is applied for classification in the second stage. One-stage methods can directly predict end-to-end bounding boxes and classification probabilities. They are faster than two-stage models, have lower computational costs, and have real-time detection capabilities.
[0054] See Figure 10 , the specific steps of a real-time detection method for rotating small targets based on multi-scale feature fusion of the present invention are as follows:
[0055] 1. Perform multi-scale feature extraction using a large receptive field multi-scale feature extraction network based on deformable convolution and dilated convolution
[0056] Infrared rotating small target detection faces multiple challenges, including small target size, diverse shapes and orientations, strong background noise interference, low contrast, etc. These factors make it difficult for traditional convolutional networks to effectively capture the key features of small targets, and they are easily interfered by background information and cannot adapt to the changing forms of rotating targets, ultimately leading to a decrease in detection accuracy. Therefore, the present invention designs a deformable dilated convolution module (DDCM), which effectively solves the problems of feature capture and expression of infrared rotating small targets by combining deformable convolution to adapt to the target morphology and multi-branch dilated convolution to expand the receptive field.
[0057] For the infrared rotating small target detection task, the present invention selects RTMDet (Real-Time Object Detectors) as the baseline method and makes improvements from multiple aspects to obtain the improved RTMDet network, which improves the detection accuracy of the model, reduces the computational cost, and enables the infrared rotating small target detector to have real-time detection capabilities. The structure of the improved RTMDet network is as Figure 1 shown, and the overall structure is divided into three main parts: the backbone network (DCNext module), the bidirectional feature fusion network (LACPAN module), and the detection head (HEAD module). This network can efficiently process object detection of multi-scale features. The input and output of each module are described in detail below.
[0058] Obtain the infrared image of the rotating small target to be detected;
[0059] In the backbone network part, the infrared image of the rotating small target to be detected is input into the backbone network. After extraction, feature maps {C_1, C_2, C_3, C_4} of different scales are output. Among them, C_1 is the first feature map, C_2 is the second feature map, C_3 is the third feature map, and C_4 is the fourth feature map. The extraction process includes the processing of multiple convolutional layers and CSP layers. On this basis, in order to further improve the receptive field and feature expression ability of the network, the present invention combines deformable convolution and dilated convolution, designs a deformable dilated convolution module (DDCM module), and processes it on the third feature map C_3 to extract a higher-level fifth feature map C_5.
[0060] In the multi-scale feature fusion stage, a bidirectional feature fusion network LACPAN is designed. By combining spatial selective large kernel convolution and adaptive rotation convolution, effective fusion of features of different scales is achieved. This design not only improves the feature expression ability but also enhances the accuracy and robustness of the network in multi-scale object detection tasks. In addition, multiple upsamplings and CSP layers are integrated in the bidirectional feature fusion network LACPAN to further enrich the expression ability of the feature maps.
[0061] In the detection head part, a shared convolution structure is adopted, but a normalization layer (BN) is independently configured for each branch. This design effectively completes the classification and regression tasks while maintaining computational efficiency, thereby improving the detection accuracy and overall performance of the model. The improved RTMDet network enhances the feature extraction ability through the DDCM module, optimizes the fusion of multi-scale features through the bidirectional feature fusion network LACPAN, and achieves efficient and accurate object detection with the detection head of shared convolution and independent normalization. These improvements enable the network to show stronger adaptability and accuracy when facing complex scenes and small object detection.
[0062] Specifically, the first part of the improved RTMDet network is the DCNeXt module, whose main function is to extract multi-scale image features from the infrared images of small rotating targets. The infrared images of small rotating targets to be detected are collected, with a size of H×W×3, where H and W are the height and width of the image respectively, and 3 is the number of channels. The infrared image is used as the input image of the DCNeXt module. The DCNeXt module includes four convolutional modules and CSP (Cross Stage Partial) layers connected to each convolutional module. By processing the input image layer by layer, four feature maps of different scales are generated, namely the first feature map C_1, the second feature map C_2, the third feature map C_3, and the fourth feature map C_4. Among them, the length and width of the first feature map C_1 are 1 / 4 of the input image, and the number of channels is 24; while the spatial resolutions of the second feature map C_2, the third feature map C_3, and the fourth feature map C_4 are halved in turn, and the number of feature channels doubles layer by layer, so as to ensure that in the deeper feature extraction, the semantic information is gradually enhanced.
[0063] Next, the third feature map C_3 is further extracted for multi-scale information through a SPPF (Spatial Pyramid Pooling Fast) layer. The SPPF layer performs multi-scale fusion on the features of the third feature map C_3 through pooling operations, enhances the global expression ability of the features while keeping the resolution unchanged, and its output feature size is the same as the resolution of the third feature map C_3, retaining the spatial information of the fourth feature map C_4. Finally, the third feature map C_3 enters the designed Deformable Dilated Convolution Module (DDCM).
[0064] The DDCM module further optimizes the context information expression ability of the third feature map C_3 by combining the characteristics of deformable convolution and dilated convolution, and at the same time improves the representation ability of high-level features. After passing through the DDCM module, the highest-level fifth feature map C_5 is generated, whose spatial resolution is the same as that of the fourth feature map C_4, which is 1 / 32 of the original image (i.e., the infrared image of the small rotating target to be detected), but the number of channels is further increased to strengthen the semantic expression.
[0065] The third part of the improved RTMDet network is the HEAD module, which is used to perform classification and regression calculations on the first multi-scale feature map P_2, the second multi-scale feature map P_3, the third multi-scale feature map P_4, and the fourth multi-scale feature map P_5 generated by the LACPAN module, so as to complete the target detection task. Specifically, in the HEAD module, each multi-scale feature map first passes through a shared convolutional layer, whose role is to standardize the feature representation so that features of different scales can be processed in the same computational framework, while reducing redundant information between channels and improving computational efficiency. The output of the shared convolutional layer is further passed to the classification and regression branches.
[0066] In the classification branch, each feature map is processed by batch normalization (BN) to reduce internal covariate shift, and then sent to the classification head for class prediction. These class prediction results pass through the Softmax function or the Sigmoid function to finally generate the class distribution of the target.
[0067] In the regression branch, after the same batch normalization process, each feature map is sent to the regression head for calculating the bounding box position of the target. The regression head usually adopts a multi-layer convolutional network to perform regression prediction on the center point position, width, height of the target, and the confidence of the target box. The classification and regression tasks share some parameters but are optimized independently to ensure the accuracy of the output results.
[0068] The multi-scale structure of the HEAD module allows each feature map to predict targets of different sizes respectively. The first multi-scale feature map P_2 is suitable for detecting smaller targets (in an image, smaller targets usually refer to targets with a relatively small aspect ratio of the object, especially those that appear relatively tiny in comparison with the background or other objects. For example, in remote sensing images, smaller targets are generally some small vehicles, pedestrians, animals, etc.), the second multi-scale feature map P_3 and the third multi-scale feature map P_4 are suitable for detecting medium-sized targets, while the fourth multi-scale feature map P_5 is mainly used for detecting larger targets. Through this division of labor, the HEAD module can make full use of the multi-scale features from the LACPAN module to achieve comprehensive detection and precise positioning of the target. Finally, the output results of classification and regression are integrated to generate the complete detection results, including the target category, location, and confidence, providing a reliable output for the target detection task.
[0069] Figure 2The structure diagram of the deformable dilated convolution module (DDCM) is shown. The input feature (i.e., the third feature map C_3) first undergoes deformable convolution, which is used to adaptively adjust the sampling positions of the convolution kernels, enhancing the ability to capture features of complex targets. Subsequently, the feature map is divided into three paths and respectively undergoes dilated convolution. Different dilated convolutions can capture multi-scale context information, effectively expanding the receptive field. After the dilated convolution, the three paths of features are fused through concatenation operations to integrate information at different scales. Finally, the fused feature map further extracts and optimizes features through depthwise separable convolution, while reducing the computational amount and improving the efficiency of the model. Finally, a residual connection is made between the initial input and the output after being processed by the above module (as shown by the "+" in the figure), obtaining the fifth feature map C_5, realizing the efficient circulation of information and alleviating the problem of gradient disappearance in the training of deep networks. This structure can effectively combine multi-scale information and spatial adaptability features, enhancing the feature expression ability of the network and the detection effect on complex targets.
[0070] In the entire backbone network, the main role of the DDCM module is to enhance the feature expression ability and receptive field of the network. The multi-scale feature maps {C_1, C_2, C_3, C_4} output by the backbone network, and the third feature map C_3 is further processed by the DDCM module to generate a higher-level fifth feature map C_5. The DDCM module adaptively adjusts the sampling positions of the convolution kernels by combining deformable convolution and dilated convolution, while capturing multi-scale context information and expanding the receptive field. The DDCM module plays a connecting role in the backbone network. It is located in the processing stage of middle and high-level feature maps, further refining and optimizing the high-dimensional semantic information of the original feature map. The generated fifth feature map C_5 has stronger expression ability and perception ability. Combined with the subsequent LACPAN module, the feature map provided by the DDCM module can better support the detection of targets at different scales, especially for small selected targets in complex backgrounds, significantly improving the accuracy and robustness of detection. Therefore, the introduction of the DDCM module in the backbone network not only enhances the effectiveness of high-level features but also provides better input features for subsequent multi-scale feature fusion and detection heads, further promoting the improvement of the overall performance of the network.
[0071] 2. Feature fusion is performed using a bidirectional feature fusion network based on spatially selective large kernel convolution and adaptive rotation convolution
[0072] When the network processes multi-scale objects and complex scenes, it faces two core problems: one is the insufficient receptive field, which cannot effectively capture global and local feature information; the other is that the features are not sensitive to the changes in the target direction, resulting in a decline in the performance of rotated object detection. To address these challenges, the present invention designs a bidirectional feature fusion network (LACPAN module) based on spatial selection large kernel convolution and adaptive rotation convolution. The LAC (Large Select Kernel Convolution, Adaptive Convolution, Cross Stage Partial Network) layer is integrated into the LACPAN network. In the LAC layer, the present invention effectively expands the receptive field and captures richer context information through spatial selection large kernel convolution. At the same time, combined with adaptive rotation convolution, the network can perceive feature changes in different directions, thereby better processing objects with rotated or irregular directions.
[0073] The second part of the improved RTMDet network is the LACPAN module, whose goal is to fuse feature maps of different scales to generate the first multi-scale feature map P_2, the second multi-scale feature map P_3, the third multi-scale feature map P_4, and the fourth multi-scale feature map P_5 for object detection, so as to better capture the context information and scale change features of the object. Specifically, first, the second feature map C_2, the third feature map C_3, the fourth feature map C_4, and the fifth feature map C_5 extracted from the DCNeXt module are sequentially input into different sub-processes of the LACPAN module. For the second feature map C_2, it is combined with the upsampling layer through concatenation operation to generate a higher-resolution ninth feature map F_2. At this time, the LAC layer will further perform local enhancement on these features, and then generate the output first multi-scale feature map P_2 after being processed by the convolution module. Next, the third feature map C_3 also undergoes concatenation and upsampling, and is combined with the upsampling layer to form the sixth feature map T_3. Then, through the LAC layer and the convolution module for further feature refinement, the intermediate-scale tenth feature map F_3 is generated. The intermediate-scale tenth feature map F_3 passes through concatenation, the CSP layer, and the convolution module, and finally outputs the second multi-scale feature map P_3. The fourth feature map C_4 passes through the convolution module to obtain the eleventh feature map F_4. The features after concatenating the ninth feature map F_2 and the tenth feature map F_3, passing through the CSP layer and the convolution module, are concatenated with the eleventh feature map F_4 and then pass through the CSP layer to obtain the third multi-scale feature map P_4.
[0074] For the fourth feature map C_4, since its features themselves have strong semantic information, it is directly combined with the CSP layer through concatenation operation for local feature optimization. The focus of this layer is to retain and enhance high-level semantic features and reduce the loss of context information. Subsequently, it is processed by the convolution module and finally outputs the third multi-scale feature map P_4.
[0075] Finally, for the fifth feature map C_5, as a deep feature map, it contains global context information and high-level semantic features of significant objects. In the LACPAN module, instead of undergoing complex fusion operations, it is directly adjusted through the channel convolution module to maintain the integrity of its global features. Finally, the fourth multi-scale feature map P_5 is output. The entire LACPAN module fully realizes the enhancement of multi-scale features through layer-by-layer feature fusion and refinement, providing high-quality input features for the subsequent HEAD module.
[0076] As Figure 3 shown, in the LAC layer, a series of convolutional operations are first performed on the input feature map (i.e., the sixth feature map T_3 or the ninth feature map F_2 with higher resolution), combined with the CSP (Cross Stage Partial) layer for feature extraction and channel decomposition, effectively improving the computational efficiency and feature representation ability of the network. The multi-level information is integrated through concatenation and convolutional operations to obtain the integrated feature map F. The integrated feature map F passes through the designed LAC layer, introducing spatially selective large kernel convolutions to capture more extensive context information, enhancing the receptive field and local expression ability of the feature map, and obtaining the seventh feature map U. Subsequently, the seventh feature map U further optimizes the information flow through the gating mechanism, screening out important feature information to obtain the eighth feature map Y. The gating mechanism extracts the global statistical information of the eighth feature map Y through global average pooling (AvgPool), and is processed through convolution and activation functions (ReLU) to generate the weights for feature selection. These weights are applied to the original feature map through element-wise multiplication, suppressing irrelevant features and highlighting key regions, and finally retaining the original feature information through residual connections to achieve effective feature screening and enhancement.
[0077] Specifically, the structural diagram of the LAC layer is as Figure 4 shown. First, the input integrated feature map F undergoes feature extraction through convolutional kernels of different sizes (1×1, 5×5, and 7×7) to capture information with different receptive fields, and then the feature map is dimension-reduced through 1×1 convolution. On this basis, two types of feature maps are extracted from the spatial dimension using average pooling and max pooling, the features are combined through channel concatenation (C) operations, and the spatial attention mechanism is adaptively adjusted through the Sigmoid activation function to obtain the weight map. At this time, the generated weight map is multiplied and added element-wise to the feature map to complete the feature fusion guided by spatial attention and obtain the fused feature map T.
[0078] Next, an adaptive rotation convolution module is introduced on the right side of the network to perform rotation perception operations on the fused feature map T, so as to improve the network's robustness to direction and rotation transformations. The output of this module is further fused with the feature map guided by left-side attention through element-wise addition operation to form a bidirectional feature stream. Finally, the fused features (i.e., the bidirectional feature stream) are subjected to channel recombination through 1×1 convolution and non-linear mapping through the SiLU activation function to complete the final output and obtain the seventh feature map U. The entire process combines multi-scale receptive fields, spatial attention mechanisms, and rotation-invariant feature extraction, effectively enhancing the network's feature expression ability for infrared selected small targets.
[0079] 3. Training Optimization Strategies Based on DSLA, MixUp, and Mosaic Models
[0080] In the task of extracting infrared rotating small targets in complex scenes, the diversity and complexity of data have an important impact on the training effect of the model. Therefore, the present invention uses data augmentation and optimization strategies of the dynamic soft label assignment algorithm (DSLA), MixUp, and Mosaic.
[0081] For the infrared rotating small target detection task, the present invention uses the Kullback-Leibler Divergence (KLD) as part of the cost matrix. KLD can dynamically adjust the parameter gradient according to the characteristics of the object. By adjusting the importance of the angle parameter according to the aspect ratio, the effectiveness of positive and negative sample assignment can be greatly improved. The present invention uses the Dynamic Soft Label Assigner (DSLA) to implement the dynamic matching strategy of labels.
[0082] The dynamic soft label assignment (DSLA) algorithm aims to achieve efficient matching of targets and anchors in the input image. Its inputs include an input image (I), a set of anchors (A), and the ground truth annotation (G) of the targets in the image. The output of the algorithm is the assignment result (π*), which is the set of anchors assigned to each target.
[0083] The specific process of the dynamic soft label assignment algorithm is as follows:
[0084] First, through the forward propagation of the network, the classification prediction (p_cls) and bounding box prediction (p_box) are obtained using the input image (I) and the anchors (A):
[0085] p_cls, p_box ← Forward(I, A).
[0086] Next, the following cost matrices are calculated in sequence:
[0087] 1. Classification cost matrix: c_cls = SoftClsCost(p_cls, G_cls), which is calculated based on the difference between the classification prediction and the target true class annotation. Here, c_cls is the classification cost matrix, and G_cls is the class label of the target.
[0088] 2. Regression cost matrix: c_iou = IoUCost(p_box, G_box), which is calculated based on the IoU (Intersection over Union) of the predicted bounding box and the ground truth bounding box. Here, c_iou is the regression cost matrix, and G_box is the location label of the target.
[0089] 3. Soft center distance matrix: c_cp = SoftCenterPrior(p_box, G_box), which is calculated from the geometric center distance between the predicted bounding box and the ground truth bounding box. Here, c_cp is the soft center distance matrix.
[0090] 4. KLD cost matrix: c_kld = KLD(p_box, G_box), which is calculated according to the distribution difference between the predicted bounding box and the ground truth bounding box. Here, c_kld is the KLD cost matrix.
[0091] Combine all cost matrices (i.e., classification cost matrix, regression cost matrix, soft center distance matrix, and KLD cost matrix) to obtain the comprehensive cost: c = c_cls + c_reg + c_cp + α * c_kld, where α is the weight parameter. For each target (G), sort the comprehensive cost in ascending order: {c1, c2,..., c10} ← AscendSort(c). Here, c1, c2, ……, c10 are the cost values of each target for each label, obtaining the sorted cost list. Determine the allocation quantity K according to the sorting result. The calculation method is K = INT(∑c). If K < 1, then take K = 1. Then, select the indices of the first K minimum cost values from the sorted cost list and add them to the allocation result π* (i.e., directly take the first K elements in Python).
[0092] The dynamic soft label assignment algorithm first calculates the position prior information cost, sample regression cost, KL divergence cost, and sample classification cost based on the input image to obtain the corresponding four cost matrices. Then, these four cost matrices are weighted and fused. Finally, SimOTA is used to determine the number of samples matched by each ground truth (GT) and decide the final samples. First, adaptively calculate the number of samples to be selected for each GT. Take the sum of the first 13 smallest costs of each GT and all predicted rotated bounding boxes, and round it up to be the number of samples for this GT, with a minimum of 1, denoted as K. After that, for each GT, take the positions of the first K smallest values in its matching cost matrix as the positive samples of this GT. Finally, for a certain predicted rotated bounding box, if it is matched to multiple GTs, take the smallest one in the matching cost matrices with these GTs as its label. The dynamic soft label assignment algorithm combines multiple cost matrices and performs dynamic matching, which can effectively improve the performance and accuracy of the object detection model. This method has different matching strategies at different stages of network training, enabling the model to achieve the best optimization effects in both the initial and later stages.
[0093] Aiming at the complex background noise in the infrared small rotated target detection task, the present invention trains the improved RTMDet network through a strong data augmentation strategy to reduce the side effects of "noisy" samples in infrared images and improve the robustness of the small rotated target detector.
[0094] The present invention uses a multi-stage data augmentation strategy to train an infrared small rotated target detector, such as Figure 5 shown. The first 10% of the iterations are in the first stage using weak data augmentation, such as random scaling and random flipping. The middle 80% of the iterations are in the second stage using strong data augmentation, including random scaling, HSV, random flipping, random rotation, Mosaic, and Mixup. The last 10% of the iterations are in the third stage, continuing to use the weak data augmentation in the first stage. In the strong data augmentation training stage, the number of mixed images in each training sample is increased to 8 to compensate for the intensity of data augmentation. In the weak data augmentation training stage, the data augmentation is weakened, and the RTMDet model is fine-tuned in a domain that is more closely aligned with the real data distribution.
[0095] Among them, when using data augmentation methods involving mixing multiple images such as Mosaic and MixUp, since the number of images to be read and processed increases exponentially, the time required for data augmentation also increases significantly. As Figure 6As shown, in order to optimize the efficiency of multi-image mixing data augmentation methods such as Mosaic and MixUp, the present invention introduces a caching mechanism. By using a cache queue to save historical images, when performing image mixing operations, instead of reloading images from the dataset, historical images are randomly selected from the cache queue for mixing. This significantly reduces the image loading time and improves the running efficiency of data augmentation. Specifically, the cached images are added to the cache queue, and when performing image mixing operations, they are randomly popped from the cache queue and randomly selected. After randomly selecting from the dataset, multiple original images are obtained.
[0096] The present invention optimizes the positive and negative sample allocation through DSLA and KLD, combines Mosaic and MixUp to enhance data diversity, and improves the efficiency through a caching mechanism, constructing an efficient model training optimization framework, and finally improving the performance and robustness of the infrared rotating small target detector.
[0097] 4. Lightweight Network Design Strategy Based on Depthwise Separable Convolution and Channel Compression
[0098] In order to improve performance and efficiency, the present invention first performs lightweight design on the network from the model structure level, adopts depthwise separable convolution and channel compression, realizes the lightweight of the network, and minimizes the computational amount and memory consumption to the greatest extent.
[0099] Depthwise separable convolution is a strategy that decomposes traditional convolution operations into two independent steps. It was first proposed by ShufflenetV2 and has been widely used in subsequent multiple lightweight networks. Depthwise separable convolution consists of two parts: depthwise convolution and pointwise convolution. Depthwise convolution applies a separate convolution kernel to each input channel for convolution operations, thus significantly reducing the computational amount and the number of parameters. Pointwise convolution then applies a 1x1 convolution to the output of the depthwise convolution to fuse information from different channels and achieve deeper feature expression.
[0100] Based on depthwise separable convolution, the present invention designs a depthwise separable convolution module (DC module) to replace the standard convolution module in RTMDet. The structure of the DC module is as Figure 8 shown, including depthwise separable convolution, batch normalization, and SiLU (Sigmoid Linear Unit) activation function. After depthwise separable convolution, it is processed by a batch normalization layer, which accelerates the convergence speed of the model and helps to stabilize the training process, preventing gradient vanishing or explosion. Subsequently, the output of batch normalization undergoes a non-linear mapping through the SiLU activation function. SiLU combines the characteristics of ReLU and Sigmoid and can better capture complex non-linear relationships in the input data, further enhancing the expression ability of the model.
[0101] Channel compression is a technique that reduces the computational load and memory consumption by decreasing the number of channels in the convolutional layers of a network. For example, reducing the output channels of a certain convolutional layer to half of the original can significantly reduce the computational load and memory consumption. As an important optimization strategy, channel compression is particularly suitable for deployment on resource-constrained devices, which can effectively improve the inference speed and reduce latency.
[0102] Based on the degree of replacement with depthwise separable convolutions and the degree of channel compression, the present invention proposes three models with different scales, namely Tiny, Micro, and Edge. Their computational loads and the number of parameters decrease in sequence. Among them, the Tiny model is the basic model in the present invention, and it achieves the best mAP (mean Average Precision) on two datasets. Depthwise separable convolutions significantly reduce the number of parameters and computational load of the network, while channel compression further reduces the network complexity by decreasing the number of channels in each layer. The present invention significantly improves the efficiency of infrared rotating small target detection by using depthwise separable convolutions on multiple layers and compressing the channels of the network.
[0103] The overall process of the real-time detection method for rotating small targets based on multi-scale feature fusion of the present invention is as Figure 9 shown. First, the input infrared image is received as the starting point of the detection process. Among them, the infrared image is subjected to feature extraction through the DCNeXt module to obtain a feature map. This module is an improved deep convolutional neural network that can effectively extract the features of small targets in the infrared image. Then, the feature map is passed to the LACPAN module, which is used to perform multi-scale feature fusion on the feature map to enhance the detection ability of rotating small targets, while improving the robustness and accuracy of the model, and obtaining the fused features. Subsequently, the HEAD module integrates and classifies the fused features to generate the target detection results. Finally, the positions of the detected small targets are output on the infrared image and labeled, with different colors representing different target categories or specific detection states.
[0104] See Figure 11 , the embodiments of the present invention also provide a real-time detection system for rotating small targets based on multi-scale feature fusion, including:
[0105] An image acquisition module for acquiring the infrared image to be detected;
[0106] A feature extraction module for performing feature extraction on the infrared image to be detected to obtain a feature map;
[0107] A multi-scale feature fusion module for performing multi-scale feature fusion on the feature map to obtain the fused features;
[0108] An integration and classification module for integrating and classifying the fused features to generate target detection results.
[0109] An embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the real-time detection method for rotated small targets based on multi-scale feature fusion is implemented.
[0110] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the real-time detection method for rotated small targets based on multi-scale feature fusion is implemented.
[0111] The actual hardware deployment of the real-time detection system for rotated small targets based on multi-scale feature fusion is as Figure 12 shown. The deployment process includes a complete detection link from system initialization to result display. First, start the program and initialize the GUI interface, including loading neural network parameters and reading the input infrared image sequence into memory. Then, extract features from the image through a convolutional neural network, and calculate the bounding boxes and categories of potential targets. The entire process uses RTX3090 hardware acceleration to improve the inference efficiency and meet the requirements of real-time detection. After the detector completes the calculation of the bounding boxes and categories, post-processing is performed through non-maximum suppression (NMS) and confidence filtering to optimize the target detection results. Finally, the processed detection results are output in coordinate form, providing accurate small target detection results for users. The entire process is simple and efficient, suitable for accurate identification and visual analysis of small targets in infrared images.
[0112] The present invention provides a new idea for solving the key technical problems of infrared rotated small target detection in complex scenarios, and has important military application prospects. The specific advantages are as follows:
[0113] The present invention aims at infrared rotated small targets and conducts real-time recognition of infrared rotated small targets in complex scenarios. Through the detection method of the present invention, a real-time detection system for infrared rotated small targets in complex scenarios will be established, which will have an important impact in future high-tech battlefields, effectively improve the combat effectiveness of the military, and play a unique and irreplaceable role.
[0114] The present invention provides a new idea for solving the key technical problems of infrared rotated small target detection in complex scenarios, and has important military application prospects.
[0115] Based on the RTMDet model, the present invention designs an efficient infrared rotating small target detection system, which is optimized for the characteristics of complex scenes. By adopting deformable convolution and dilated convolution, a deformable dilated convolution module (DDCM) is designed to enhance the feature extraction ability of the backbone network. At the same time, based on spatial selective large kernel convolution and adaptive rotation convolution, a bidirectional feature fusion network LACPAN is designed to effectively combine features of different scales and enhance the feature expression of infrared rotating small targets. To improve the performance of the model in complex environments, a dynamic soft label assignment algorithm and a multi-stage data augmentation strategy (such as Mosaic, MixUp and random rotation) are adopted to make the training more sufficient and achieve accurate recognition of infrared rotating small targets in complex scenes. Aiming at the high complexity of the infrared rotating small target detection model, the present invention completes the lightweight research of the rotating small target detection model in complex scenes through an optimized process. A lightweight network is designed by using depthwise separable convolution and channel compression technology, which reduces the computational amount while maintaining the detection accuracy.
[0116] The successful implementation of the present invention not only improves the technical level of infrared rotating small target detection, but also provides strong technical support for the field of intelligent combat analysis and processing. Through the real-time detection system, high-quality video image data can be effectively analyzed, different types of targets can be identified and their positions can be accurately located, a multi-dimensional information situation awareness system can be established, the digitization and networking of battlefield information can be realized, and the battlefield dynamics can be comprehensively and accurately grasped. In the military field, infrared small target detection technology can be used to scout and monitor distant targets, such as unmanned aerial vehicles, missiles, etc.; in the security field, infrared small target detection technology can timely detect abnormal targets and improve public security. This will greatly enhance the intelligence reconnaissance, surveillance tracking and precision strike capabilities of unmanned intelligent weapons such as unmanned aerial vehicles, save a large amount of manpower and time costs, and ensure the initiative in war.
[0117] In addition, with the rapid development of unmanned aerial vehicle remote sensing infrared image technology, infrared imaging device reconnaissance has become an important means to perceive enemy battlefield information. The research results of the present invention will effectively enhance the combat effectiveness of the military, play a unique and irreplaceable role, and have important significance for improving the national security defense ability and coping with the complex and changeable security environment.
[0118] The present invention not only makes a breakthrough in technology, but also has broad application prospects in military and civilian fields. With the further maturity and improvement of technology, it is expected to play a more important role in future intelligent construction and security monitoring.
[0119] The above description is only for the best embodiments of the present invention, but it should not be construed as a limitation on the claims. The present invention is not limited to the above embodiments, and its specific structure allows for variations. Any variations made within the scope of protection of the independent claims of the present invention are within the scope of protection of the present invention.
[0120] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
Claims
1. A real-time detection method for rotating small targets based on multi-scale feature fusion, characterized in that: The following steps are involved: Acquire an infrared image to be detected; Perform feature extraction on the infrared image to be detected to obtain a feature map; Perform multi-scale feature fusion on the feature map to obtain fused features; The fused features are integrated and classified to generate target detection results.
2. The real-time detection method for rotating small targets based on multi-scale feature fusion according to claim 1 is characterized in that: Performing feature extraction on the infrared image to be detected to obtain a feature map includes the following steps: The infrared image to be detected is subjected to feature extraction to obtain feature maps of different scales {C_1, C_2, C_3, C_4}, where C_1 is the first feature map, C_2 is the second feature map, C_3 is the third feature map, and C_4 is the fourth feature map. The extraction process includes multi-layer convolution and CSP layer processing; the third feature map C_3 is subjected to deformable convolution and hole convolution to obtain the fifth feature map C_5.
3. The real-time detection method for rotating small targets based on multi-scale feature fusion according to claim 2 is characterized in that: The length and width of the first feature map C_1 are 1 / 4 of the infrared image to be detected, and the number of channels is 24; the spatial resolutions of the second feature map C_2, the third feature map C_3, and the fourth feature map C_4 are halved successively, and the number of feature channels is doubled layer by layer.
4. The real-time detection method for rotating small targets based on multi-scale feature fusion according to claim 1 is characterized in that: The third feature map C_3 is subjected to deformable convolution and dilated convolution to obtain the fifth feature map C_5, including the following steps: The feature map of the third feature map C_3 after deformable convolution is divided into three paths, which are respectively subjected to hollow convolution, and then the three-path features are fused through a series operation. The fused feature map is extracted and optimized through depthwise separable convolution. The third feature map C_3 and the output after feature extraction and optimization through depthwise separable convolution are residually connected to obtain the fifth feature map C_5.
5. The real-time detection method for rotating small targets based on multi-scale feature fusion according to claim 2 is characterized in that: Performing multi-scale feature fusion on the feature map to obtain fused features includes the following steps: The second feature map C_2 is connected to the upsampling layer in series to generate the ninth feature map F_2, the ninth feature map F_2 is locally enhanced, and then processed by the convolution module to generate the first multi-scale feature map P_2; The third feature map C_3 is connected in series and processed by the upsampling layer to form the sixth feature map T_3; the sixth feature map T_3 is refined by the LAC layer and the convolution module to generate the tenth feature map F_3; the tenth feature map F_3 is processed by the connection, CSP layer and convolution module to generate the second multi-scale feature map P_3; The fourth feature map C_4 is processed through the convolution module to obtain the eleventh feature map F_4. The features of the ninth feature map F_2 and the tenth feature map F_3 are concatenated through the CSP layer and the features of the convolution module. After being concatenated with the eleventh feature map F_4 through the CSP layer, the third multi-scale feature map P_4 is obtained. The fifth feature map C_5 is adjusted through the channel convolution module to generate the fourth multi-scale feature map P_5.
6. The real-time detection method for rotating small targets based on multi-scale feature fusion according to claim 5 is characterized in that: In the LAC layer, the sixth feature map T_3 or the ninth feature map F_2 is firstly convolved, and the CSP layer is combined to perform feature extraction and channel decomposition, and then the multi-level information is integrated through concatenation and convolution operations to obtain the integrated feature map F; the integrated feature map F passes through the LAC layer to obtain the seventh feature map U; the seventh feature map U optimizes the information flow through the gating mechanism to obtain the eighth feature map Y; wherein, the gating mechanism extracts the global statistical information of the features through global average pooling, and processes it through convolution and activation functions to generate feature selection weights; the feature selection weights are applied to the seventh feature map U through element-by-element multiplication, and finally the original feature information is retained through the residual connection to obtain the eighth feature map Y.
7. The real-time detection method for rotating small targets based on multi-scale feature fusion according to claim 1 is characterized in that: Integrate and classify the fused features to generate target detection results, including the following steps: The first multi-scale feature map P_2, the second multi-scale feature map P_3, the third multi-scale feature map P_4, and the fourth multi-scale feature map P_5 are passed to the classification branch and the regression branch after passing through the shared convolution layer; In the classification branch, each feature map is processed by batch normalization and then sent to the classification head for category prediction. The category prediction result is passed through the Softmax function or the Sigmoid function to generate the category distribution of the target; In the regression branch, each feature map is batch normalized and then fed into the regression head to calculate the bounding box position of the target.
8. A real-time detection system for rotating small targets based on multi-scale feature fusion, characterized in that: include: An image acquisition module, used for acquiring an infrared image to be detected; A feature extraction module is used to extract features from the infrared image to be detected and obtain a feature map; A multi-scale feature fusion module is used to perform multi-scale feature fusion on feature maps to obtain fused features; The integration and classification module is used to integrate and classify the fused features to generate target detection results.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the real-time detection method for rotating small targets based on multi-scale feature fusion as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the real-time detection method for rotating small targets based on multi-scale feature fusion as described in any one of claims 1 to 7 is implemented.