Transform model for real-time image detection and application thereof

Through the Transformer model with a bidirectional receptive field optimization module, a deformable attention mechanism, and an attention upsampling module, the problems of complex background and weak features of small targets in SAR images are solved, and high-precision multi-scale target detection is achieved.

CN120708081AActive Publication Date: 2025-09-26ANHUI UNIV

Patent Information

Application Number
CN202510803697.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The problem of reduced recognition accuracy in SAR image target detection due to complex background interference, weak small target features and coexistence of multi-scale targets.

Method used

The Transformer model, which adopts a bidirectional receptive field optimization module, a deformable attention mechanism module, and an attention upsampling module, improves the target feature expression capability by efficiently fusing local details with global context information, dynamically adjusting the feature sampling point position, and multi-path feature processing.

Benefits of technology

The accuracy and robustness of SAR image target detection are significantly improved, especially the recognition ability of small targets in complex scenes, reducing missed detections and false detections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708081A_ABST
    Figure CN120708081A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and remote sensing, and particularly relates to a Transform model for real-time image detection and application thereof, and the model mainly comprises a bidirectional receptive field optimization module, a deformable attention mechanism module and an attention up-sampling module. The bidirectional receptive field optimization module enhances the key feature expression ability in a mode of combining feature extraction and an attention mechanism; the deformable attention mechanism module adopts a dynamic sampling strategy to adaptively focus a target area; the attention up-sampling module maintains detail information through multi-path feature fusion. The model effectively solves the problems of complex background interference, weak small target features, multi-scale target detection and the like in the SAR image, remarkably improves the detection precision and robustness, and can be widely applied to the remote sensing fields of military reconnaissance, ocean monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and remote sensing technology, and particularly relates to a Transformer model for real-time image detection and its application. Background Art

[0002] SAR imagery, with its cloud-penetrating and all-weather observation capabilities, has important applications in ocean monitoring, military reconnaissance, and other fields. SAR image object detection technology has long been a hot research topic in this field. Early statistical modeling-based constant false alarm rate (CFAR) algorithms relied on artificial feature design, limited generalization, and struggled to adapt to complex backgrounds and multi-scale objects. In recent years, deep learning technology has made significant progress in SAR image object detection. Researchers have primarily focused on multi-scale modeling and dynamic attention mechanisms to improve detection accuracy and robustness. For example, two-stage object detection algorithms such as FasterR-CNN and MaskR-CNN utilize region proposal networks and feature pyramid networks to achieve end-to-end detection, significantly improving both accuracy and speed. Furthermore, single-stage object detection algorithms such as YOLOv5 and DETR further enhance detection speed and efficiency through the use of anchor box mechanisms and the Transformer architecture. The application of attention mechanisms such as CBAM and the Swin Transformer has also effectively enhanced the model's ability to focus on key features and mitigate noise interference.

[0003] SAR images are susceptible to natural environmental noise contamination, such as clouds, shadows, rain, and fog. This can blur target edges and lose texture information, making it difficult for models to accurately identify targets. Furthermore, due to the limitations of SAR imaging resolution, distant targets occupy a very small portion of the image, and their feature information is sparse, making them easily overlooked or misjudged by the algorithm. Furthermore, the large variations in target scale in SAR images make it difficult for traditional detection algorithms to adaptively capture features at different scales, leading to missed or false detections. These issues collectively limit the accuracy and robustness of SAR detection algorithms, especially in complex scenarios. Summary of the Invention

[0004] The purpose of the present invention is to provide a Transformer model for real-time image detection and its application to solve the problem of decreased recognition accuracy in SAR image target detection due to complex background interference, weak small target features and coexistence of multi-scale targets.

[0005] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0006] This paper proposes a Transformer model for real-time image detection, including:

[0007] The bidirectional receptive field optimization module is used to extract the spatial features of the input image, perform dimensionality conversion to form a first feature map, multiply the first feature map with the attention map, generate a second feature map through grouped convolution and normalization operations, and perform a residual connection between the second feature map and the input image to generate a first feature map to be processed;

[0008] The deformable attention mechanism module is used to generate N uniform grid points as reference points in the first feature map to be processed, predict the offset of the reference points and obtain the position of the deformation points through a lightweight sub-network, use bilinear interpolation to sample features at the deformation point positions, generate keys and values, and form the second feature map to be processed;

[0009] The attention upsampling module is used to divide the second feature map to be processed into three processing paths. The first path improves the resolution of the feature map through the transposed convolution operation. The second path combines upsampling with convolution to compensate and refine the detail information in the feature map. The third path uses global pooling to extract global features and calibrates the feature weights through the channel attention mechanism. The outputs of the first and second paths are spliced ​​and multiplied with the weights of the third path, and then the target feature map is output through convolution.

[0010] Preferably, the bidirectional receptive field optimization module includes:

[0011] The first convolution unit is used to extract spatial features through a k×k convolution kernel and convert the input dimension from C×H×W to a first feature map of C′×H×W;

[0012] A feature adjustment unit, configured to perform matrix multiplication of the first feature map and the attention map;

[0013] A second convolution unit is used to perform grouped convolution and normalization on the multiplication result, and output a second feature map with a dimension of C×H×W;

[0014] The residual connection unit is used to add the input feature map to the output of the second convolutional unit to form a forward-feedback feature interaction mechanism and generate a first feature map to be processed with a dimension of 2C×H×W.

[0015] Preferably, the deformable attention mechanism module includes:

[0016] A reference point generating unit, configured to generate N evenly distributed grid points in the first feature map to be processed as reference points, wherein the grid size is determined based on a downsampling factor;

[0017] The offset prediction unit is used to predict the offset of each reference point through a lightweight sub-network and scale the offset by a predefined factor to obtain the position of the deformation point;

[0018] A feature sampling unit is used to sample features at the deformation point using a bilinear interpolation function to generate keys and values;

[0019] The multi-head attention unit is used to perform multi-head attention calculation on the sampled features and introduce relative position offset to form a second feature map to be processed.

[0020] Preferably, in the reference point generating unit:

[0021] The reference points are linearly spaced two-dimensional coordinates normalized to the range of [-1, 1];

[0022] The grid size satisfies the formula s=H / d×W / d, where d is the downsampling factor, and H and W are the feature map height and width, respectively.

[0023] Preferably, in the multi-head attention unit, multi-head attention calculation is performed on the sampled features, and relative position offset is introduced, including:

[0024] Set up multiple attention heads, each of which independently predicts deformation offset and attention weight;

[0025] Introduce locality constraints to limit the sampling range of each reference point to a 3×3 neighborhood;

[0026] Absolute position encoding and relative position deviation table are integrated in attention calculation.

[0027] Preferably, in the attention upsampling module:

[0028] The first processing uses 1×1 transposed convolution to improve feature resolution;

[0029] The second path of processing sequentially performs an upsampling operation and a 1×1 convolution operation;

[0030] The third processing step sequentially performs global pooling, 1×1 convolution, and Hardsigmoid activation function processing.

[0031] Preferably, a hybrid aggregation network module is also included to enhance multi-scale feature fusion by:

[0032] Deploy a multi-level feature pyramid structure to integrate semantic information at different scales;

[0033] Use cross-layer skip connections to reduce the feature degradation problem of deep networks;

[0034] The gradient propagation efficiency of the feature pyramid is optimized through an adaptive weight distribution mechanism.

[0035] Preferably, in the hybrid aggregation network module, the cross-layer jump connection is specifically:

[0036] The shallow feature map is aligned with the channel dimension through 1×1 convolution and then added element-by-element to the deep feature map;

[0037] The weights of the skip connections are dynamically adjusted through a learnable gating mechanism to suppress redundant feature interference;

[0038] The fused feature map undergoes 3×3 convolution to further optimize the spatial information expression.

[0039] The present invention also proposes an application of the above-mentioned Transformer model for real-time image detection in SAR image target detection. When the model is applied to SAR image target detection, the method includes the following steps:

[0040] S1. Input the SAR image to the bidirectional receptive field optimization module, extract spatial features using a k×k convolution kernel, and use a residual connection unit to suppress sea clutter interference, outputting a first feature map to be processed with a dimension of 2C×H×W;

[0041] S2. Input the first feature map to be processed into the deformable attention mechanism module, generate a reference point normalized to the range of [-1, 1], focus on the scattering area of ​​the ship, aircraft or vehicle target through bilinear interpolation sampling, and output a dynamically adjusted second feature map to be processed;

[0042] S3. Perform three-way processing on the dynamically adjusted second feature map to be processed by the attention upsampling module;

[0043] S4. Input the target feature map processed in step S3 into the hybrid aggregation network module, fuse the multi-scale features through cross-layer skip connections, and output the final detection result, which includes the target category and bounding box coordinates.

[0044] The beneficial effects of the present invention are:

[0045] The real-time detection Transformer model based on SAR complex background provided by the present invention has significant technical advantages and practical value. The model realizes the efficient fusion of local detail features and global context information through the bidirectional receptive field optimization module, effectively overcomes the defect of insufficient feature extraction ability of traditional methods under complex backgrounds, and significantly improves the expression ability of key features of the target. The deformable attention mechanism module enables the network to adaptively focus on the effective scattering area of ​​the target by dynamically adjusting the position of feature sampling points, especially solving the problem that small targets have weak features and are easily submerged by the background. The attention upsampling module adopts a multi-path parallel processing architecture, and through a combination of transposed convolution, upsampling compensation and channel attention weighting, it maximizes the retention of target detail features and improves the reconstruction quality of edge information. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a structural diagram of the Transformer model for real-time image detection proposed in the present invention;

[0047] Figure 2 Schematic diagram of the structure of the BRFbo module in an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the structure of the DeformableAIFI module in an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram of the structure of the AttUp module in an embodiment of the present invention;

[0050] Figure 5 Schematic diagram of the structure of the SA-RTDETR model in an embodiment of the present invention;

[0051] Figure 6 This is a schematic diagram of the application process of the Transformer model of real-time image detection in SAR image target detection of the present invention;

[0052] Figure 7 1 is a comparison chart of ship target detection results in an embodiment of the present invention;

[0053] Figure 8 1 is a comparison chart of aircraft target detection results in an embodiment of the present invention;

[0054] Figure 9 2 is a comparison chart of the automobile target detection results in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The following description provides specific application scenarios and requirements for this specification, with the goal of enabling those skilled in the art to make and use the contents of this specification. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but is intended to be accorded the broadest scope consistent with the claims.

[0056] The terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. For example, as used herein, the singular forms "a," "an," and "the" may also include the plural forms unless the context clearly indicates otherwise. When used in this specification, the terms "comprise," "include," and / or "contain" are intended to refer to the presence of the associated integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.

[0057] These and other features of this specification, as well as the operation and function of the associated elements of the structure, and the economical assembly and manufacture of the components, can be significantly improved with consideration of the following description. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be expressly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0058] The flowcharts used in this specification illustrate operations implemented by systems according to some embodiments of the present specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. Rather, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0059] The present invention designs a real-time detection model based on SAR complex targets, aiming to enhance the ability to extract target features in order to improve the accuracy of detection in SAR images. The model is named Synthetic Aperture-optimized Real-Time Detection Transformer (SA-RTDETR) model based on synthetic aperture radar image optimization. The model fuses local details with global context information through the bidirectional receptive field boosting (BRFbo) module, significantly improving the key feature extraction capability while maintaining feature resolution. The deformable attention mechanism (Deformable Attention-based Intra-scale Feature Interaction, Deformable AIFI) module adaptively samples key feature points to enable the network to focus on the effective scattering area of ​​small targets. The attention upsampling (AttUp) module solves the problems of detail loss and image distortion through feature compensation strategies.

[0060] Example 1

[0061] See also Figure 1-5 , proposed a Transformer model for real-time image detection, including:

[0062] The bidirectional receptive field optimization module (BRFbo module) is used to extract the spatial features of the input image, perform dimensionality conversion to form a first feature map, multiply the first feature map with the attention map, and then generate a second feature map through grouped convolution and normalization operations. The second feature map is residually connected with the input image to generate the first feature map to be processed;

[0063] The Deformable AIFI module is used to generate N uniform grid points in the first feature map to be processed as reference points, predict the offset of the reference points and obtain the positions of the deformation points through a lightweight sub-network, sample features at the deformation points using bilinear interpolation, generate keys and values, and form the second feature map to be processed;

[0064] As you can understand, in the deformable attention module, sampling results are calculated using a multi-head attention mechanism, where each attention head independently predicts deformation offsets and weights, while incorporating a locality constraint within a 3×3 neighborhood. The attention calculation integrates the absolute position encoding and the relative position deviation table, ultimately outputting a dynamically adjusted second feature map to be processed. This mechanism significantly improves the localization accuracy of small objects, such as ship wingtips.

[0065] The attention upsampling module (AttUp module) is used to divide the second feature map to be processed into three processing paths: the first path improves the resolution of the feature map through the transposed convolution operation; the second path further compensates and refines the detail information in the feature map by combining upsampling and convolution; the third path uses global pooling to extract global features and calibrates the feature weights through the channel attention mechanism; the outputs of the first and second paths are spliced ​​and multiplied with the weights of the third path, and then convolved to output the final feature map.

[0066] More specifically, the Bidirectional Receptive Field Optimization module (BRFbo module) achieves efficient fusion of local details and global context information through a dual-path collaborative working mechanism. This module further includes:

[0067] The first convolution unit is used to extract spatial features through the k×k convolution kernel and convert the input dimension from C×H×W to the first feature map of C′×H×W; the formula is as follows:

[0068] FConv=ReLU(Norm(g k×k (X)))

[0069] Where g k×kRepresents a convolution with a kernel size of k×k, Norm represents normalization, and ReLU represents the ReLU activation function.

[0070] The feature adjustment unit is used to perform matrix multiplication of the first feature map output by the first convolution unit and the attention map. The formula is as follows:

[0071] TConv=FConv×Softmax(g 1×1 (AvgPool(X)))

[0072] A second convolution unit is used to perform grouped convolution and normalization on the multiplication result, and output a second feature map with a dimension of C×H×W;

[0073] SConv=Conv(TConv)

[0074] The residual connection unit is used to add the input feature map to the output of the second convolutional unit to form a forward-feedback feature interaction mechanism, and finally generate a first feature map to be processed with a dimension of 2C×H×W.

[0075] BRFbo=SConv+X

[0076] When the solution is implemented, the working process of the bidirectional receptive field optimization module is as follows:

[0077] First, the input SAR image (of dimensions C×H×W) is processed by the first convolutional unit. This unit uses a k×k convolution kernel (preferably k=3) for spatial feature extraction. The convolution operation expands the number of input channels from C to C′ (typically C′=2C), while maintaining the feature map size HxW. The convolution is followed by batch normalization (BatchNorm) and a ReLU activation function. This step effectively captures multi-scale contextual information, particularly enhancing the edge texture of ship targets and the highlights of aircraft tail flames in SAR images.

[0078] Next, the feature adjustment unit performs matrix multiplication on the first feature map and the attention map. The attention map is generated through an additional branch network, with the same dimensions as the first feature map (C' × H × W). This multiplication operation essentially implements spatial weighting of features, enabling the network to adaptively focus on the target's effective scattering area (such as a ship's deck or vehicle taillights) while suppressing background interference such as sea clutter.

[0079] The second convolutional unit then performs group convolution and normalization on the weighted feature map. Group convolution uses a group size of G = 8, reducing the number of parameters while maintaining feature interaction. The output feature map is restored to its original dimensions of C × H × W.

[0080] Finally, the residual connection unit adds the original input feature map to the output of the second convolutional unit, forming a forward-feedback feature interaction mechanism. This design has two advantages: first, it alleviates the vanishing gradient problem of deep networks through skip connections; second, it preserves the underlying scattering properties of the original SAR image (such as the strong reflective areas of ships). The output feature map is expanded to 2C×H×W dimensions.

[0081] More specifically, the deformable attention mechanism module further includes:

[0082] A reference point generation unit, configured to generate N evenly distributed grid points in the first feature map to be processed as reference points, where the grid size is determined by the downsampling factor;

[0083] In practice, the input feature map (dimensions 2C × H × W) is first divided into a regular grid. The grid size is determined by the downsampling factor d (typical value d = 16), calculated as s = H / d × W / d. Each grid point generates linearly spaced two-dimensional coordinates (e.g., a 5×5 grid corresponds to 25 reference points), and the coordinate values ​​are normalized to the range [-1, 1]. This design allows the reference points to cover target areas of varying scales, and is particularly adaptable to ship targets (occupying 0.5%-5% of the image area).

[0084] The offset prediction unit is used to predict the offset of each reference point through a lightweight sub-network, and scale the offset by a predefined factor to obtain the position of the deformation point;

[0085] In specific implementation, the feature map generates query tokens through linear projection Where N is the number of reference points and C_q is the number of query channels (default 256). The lightweight subnetwork (composed of two layers of 1×1 convolution) predicts the offset of each reference point based on Q To avoid training instability, the offset is scaled by a predefined factor λ (empirical value λ = 0.1).

[0086] A feature sampling unit is used to sample features at the deformation point using a bilinear interpolation function to generate keys and values;

[0087] The multi-head attention unit is used to perform multi-head attention calculation on the sampled features and introduce relative position offset to form a second feature map to be processed.

[0088] More specifically, in the reference point generation unit: the reference point values ​​are linearly spaced two-dimensional coordinates, normalized to the range of [-1, 1]; the grid size satisfies the formula s = H / d × W / d, where d is the downsampling factor, H and W are the feature map height and width, respectively.

[0089] More specifically, in the offset prediction unit: the feature map is linearly projected to the query token q = xW q On; input the query token into a lightweight sub-network θ offset (·) to generate the offset Δp = θ offset (q); define the scaling factor s to scale the amplitude of Δp to obtain the position of the deformation point, that is, Δp←stanh(Δp).

[0090] More specifically, in the feature sampling unit: let the sampling function φ(·; ·) be a differentiable bilinear interpolation function, that is:

[0091]

[0092] where g(a, b) = max(0, 1-|ab|), (r x , r y )express All locations on .

[0093] The sampling results are as keys and values, i.e.:

[0094]

[0095] in and denote the keys and values ​​embedded after deformation, respectively, and φ(·; ·) is the sampling function.

[0096] More specifically, in the multi-head attention unit, multi-head attention calculation is performed on the sampled features, and relative position offset is introduced, including: setting multiple attention heads, each attention head independently predicts the deformation offset and attention weight; introducing local constraints to limit the sampling range of each reference point to a 3×3 neighborhood; and fusing absolute position encoding and relative position deviation table in the attention calculation.

[0097] More specifically, in the attention upsampling module:

[0098] The first processing uses 1×1 transposed convolution to improve feature resolution;

[0099] Y1=ConvTranspose(X)

[0100] The second processing path performs upsampling and 1×1 convolution operations in sequence;

[0101] Y2=Conv(Upsample(X))

[0102] The third processing step performs global pooling, 1×1 convolution and Hardsigmoid activation function processing in sequence.

[0103] Y3=Hardsigmoid(Conv(Globalpool(X)))

[0104] The feature weights are calibrated through the channel attention mechanism, the outputs of the first and second paths are concatenated and multiplied with the weights of the third path, and then the target feature map is output through convolution.

[0105]

[0106] More specifically, it also includes a hybrid aggregation network module, which enhances multi-scale feature fusion through the following methods: deploying a multi-level feature pyramid structure to fuse semantic information of different scales; using cross-layer skip connections to reduce the feature degradation problem of deep networks; and optimizing the gradient propagation efficiency of the feature pyramid through an adaptive weight allocation mechanism.

[0107] In this embodiment, the above modules are combined with a hybrid aggregation network (MANet) to form a complete model. The input SAR image is sequentially processed through the BRFbo module to suppress noise, the Deformable AIFI module to dynamically focus on the target area, and the AttUp module to reconstruct details. Finally, multi-scale features are fused through cross-layer skip connections to output the target category and bounding box.

[0108] More specifically, in the hybrid aggregation network module, the cross-layer jump connection is as follows: the shallow feature map is aligned with the channel dimension through 1×1 convolution and then added element-by-element with the deep feature map; the weight of the jump connection is dynamically adjusted through a learnable gating mechanism to suppress redundant feature interference; the fused feature map is further optimized through 3×3 convolution to further optimize the spatial information expression.

[0109] Example 2

[0110] See also Figure 6 This embodiment proposes an application of the real-time image detection Transformer model in Example 1 to SAR image target detection. When the model is applied to SAR image target detection, the following steps are included:

[0111] S1. Input the SAR image to the bidirectional receptive field optimization module, extract spatial features using k×k convolution kernels, and use residual connection units to suppress sea clutter interference, outputting an optimized feature map with dimensions C×H×W;

[0112] S2. Input the optimized feature map into the deformable attention mechanism module, generate reference points normalized to the range [-1, 1], focus on the scattering area of ​​the ship, aircraft, or vehicle target through bilinear interpolation sampling, and output the dynamically adjusted feature map;

[0113] S3. Perform three-way processing on the dynamically adjusted feature map through the attention upsampling module;

[0114] S4. Input the feature map processed in step S3 into the hybrid aggregation network module, fuse the multi-scale features through cross-layer skip connections, and output the final detection result, which includes the target category and bounding box coordinates.

[0115] It can be understood that the above Transformer model is suitable for at least one of the following application scenarios: target recognition in military reconnaissance; ship detection in ocean monitoring; all-weather remote sensing image analysis.

[0116] After completing the above module design and combining it with the benchmark model, a real-time monitoring Transformer model based on SAR complex background can be obtained.

[0117] In this embodiment, a neural network model is built under the PyTorch framework. Figure 5 FIG. 1 shows a schematic diagram of the SA-RTDETR model network architecture in an embodiment of the present invention. Figure 5 As shown in the figure, the SA-RTDETR model takes SAR images as input. It first extracts features using cascaded 3×3 convolutional blocks (3×3Conv) and max pooling (Maxpool), and preserves large-scale information through the bidirectional receptive field optimization (BRFbo) module. It then uses a deformable attention-based intra-scale feature interaction mechanism (Deformable AIFI) to dynamically adjust the important regions of target features in the SAR image. The attention upsampling (AttUp) module, combined with a hybrid aggregation network (MANet), reduces feature information loss, improves the baseline model's ability to perceive small targets in complex backgrounds, and enhances the model's recognition accuracy. Finally, an Intersection-over-Union (IoU)-aware Query Selection mechanism is introduced at the decoder end to significantly reduce the model's parameter size while maintaining detection accuracy. Combining these modules, the model generates target predictions through the decoder head.

[0118] The above scheme is explained and analyzed in conjunction with a specific simulation experiment.

[0119] In order to compare the performance of the SA-RTDETR model, seven other models were built: Faster R-CNN, RetinaNet, FCOS, YOLOv12-L, DETR, Swin Transformer and RT-DETR50.

[0120] Simulation Experiment 1

[0121] The model provided by the present invention is used to identify three types of targets: ships, aircraft, and vehicles. The detection performance of SA-RTDETR is verified by comparing the detection results of different models.

[0122] During the training process, the input image size was uniformly adjusted to 256×256; the AdamW optimizer was used, the initial learning rate was set to 1×10-4, and the training cycle was 300; non-maximum suppression was used to filter candidate boxes, and the intersection-over-union ratio threshold was set to 0.5.

[0123] Figure 7 The results of the ship target comparison test in the embodiment of the present invention are shown in FIG. In a complex nearshore scene containing 6 ship targets, Figure 7 The SA-RTDETR algorithm shown in (b) shows a significant advantage, with all ships being accurately detected (confidence 0.75-0.89). Figure 7 (g) shows the detection results of the YOLOv12-L algorithm. The forward-feedback feature interaction mechanism formed by the two-layer convolutional blocks of the BRFbo module designed by this algorithm effectively suppresses the interference of wave clutter, and realizes feature calibration through residual connection fusion, accurately identifying the ship in the lower left corner and avoiding misidentification. It is important to note that Figure 7 The RT-DETR50 model detection results shown in (c) show that its global attention mechanism, which lacks a spatial offset, results in the loss of local features, misidentifying wave texture as a ship. The AttUp module significantly improves the jagged edges of the ship by fusing three-way sampling. In summary, this model accurately identifies all ship models in this detection task, validating its reliability in identifying small objects in complex backgrounds.

[0124] Figure 8 The results of the aircraft target comparison test in the embodiment of the present invention are shown in FIG. SA-RTDETR accurately detects all targets in the aircraft target detection test with a probability greater than 0.9, and the detection accuracy is significantly improved. This is mainly due to the double-layer residual feedback mechanism of the BRFbo module, which can enhance the dark background noise suppression. Figure 8 In (b), the aircraft tail plume area is accurately distinguished, while Figure 8 (d) and Figure 8 (e) High-brightness clouds are mistakenly identified as aircraft. Furthermore, the dynamic attention mechanism of the DeformableAIFI module accurately locates small-scale features on the aircraft's wingtips, while the three-way upsampling strategy of the AttUp module effectively preserves the jagged edges of the aircraft. These data demonstrate that the model's fusion of multi-scale upsampling and multi-branch convolution effectively reduces the blurring of target edge textures under extremely dark backgrounds, validating the effectiveness of its multi-module collaborative optimization.

[0125] Figure 9The figure shows the results of the vehicle target comparison test in the embodiment of the present invention. SA-RTDETR accurately detects 4 vehicle targets in a dark environment with a probability close to 0.9, which is significantly better than other models. Figure 9 The vehicle taillight area in (b) is accurately distinguished. This is due to the BRFbo module's two-layer convolutional block, which forms a forward-feedback feature interaction mechanism that effectively suppresses dark background noise. This avoids the misdetection of targets as ships, as occurs with DETR. Furthermore, the FCOS model also accurately detects the target, but its detection probability is lower than that of our model. Overall, the noise suppression of the BRFbo module and the edge reconstruction capabilities of the AttUp module significantly improve our model's detection accuracy for low-contrast targets.

[0126] As shown in Table 1, SA-RTDETR demonstrates significant advantages in the global performance comparison, achieving the highest recall (R = 84.7%), mAP50 (90.1%), and mAP50-95 (56.0%) among all models, representing improvements of 2.2, 2.7, and 2.6 percentage points, respectively, compared to the RT-DETR50 model (R = 82.5%, mAP50 = 87.4%, and mAP50-95 = 53.4%). Test data demonstrates that the BRFbo module, through its convolutional block feature extraction, the dynamic focusing capabilities of Deformable AIFI, and the three-way feature reconstruction capabilities of AttUp, effectively suppresses complex background interference and reduces missed detections. Although its recall is slightly lower than that of the FCOS model, it exhibits significant advantages over the FCOS model in all other metrics. This demonstrates that the SA-RTDETR model achieves the best performance, with high accuracy and strong generalization.

[0127] Table 1

[0128]

[0129] Compared with other models, the SA-RTDETR model achieved the highest average accuracy across various target detection results, with fewer missed and false detections and no significant detection errors. This demonstrates that the SA-RTDETR model exhibits high prediction accuracy across all target types, further demonstrating its reliability and stability.

[0130] In summary, the SA-RTDETR model proposed in this embodiment of the present invention uses the BRFbo module to capture key features in SAR images, the DeormableAIFI attention mechanism to adjust the important regions of target features in SAR images, and the AttUp attention upsampling operator to reduce information loss from upsampling, thereby improving the multi-scale recognition capabilities of SAR images. The model's reliability and prediction accuracy were verified by comparing it with RTDETR-50, FasterR-CNN, RetinaNet, FCOS, YOLOv12-L, DETR, and SwinTransformer. A case study was also conducted on three typical target categories: ships, aircraft, and vehicles. The results show that SA-RTDETR achieves a mAP50 of 90.1%, a mAP50-95 of 56.0%, and an improved recall (R) of 84.7%. Compared to other deep learning models, it offers higher detection accuracy and precision, avoiding the problems of missed and high false positives in multi-target SAR images, and provides a new technical paradigm for all-weather and all-day remote sensing target detection.

[0131] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0132] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0133] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.

Claims

1. A Transformer model for real-time image detection, characterized in that: include: The bidirectional receptive field optimization module is used to extract the spatial features of the input image, perform dimensionality conversion to form a first feature map, multiply the first feature map with the attention map, generate a second feature map through grouped convolution and normalization operations, and perform a residual connection between the second feature map and the input image to generate a first feature map to be processed; The deformable attention mechanism module is used to generate N uniform grid points as reference points in the first feature map to be processed, predict the offset of the reference points and obtain the position of the deformation points through a lightweight sub-network, use bilinear interpolation to sample features at the deformation point positions, generate keys and values, and form the second feature map to be processed; The attention upsampling module is used to divide the second feature map to be processed into three processing paths. The first path improves the resolution of the feature map through the transposed convolution operation. The second path combines upsampling with convolution to compensate and refine the detail information in the feature map. The third path uses global pooling to extract global features and calibrates the feature weights through the channel attention mechanism. The outputs of the first and second paths are spliced ​​and multiplied with the weights of the third path, and then the target feature map is output through convolution.

2. A Transformer model for real-time image detection according to claim 1, characterized in that: The bidirectional receptive field optimization module includes: The first convolution unit is used to extract spatial features through a k×k convolution kernel and convert the input dimension from C×H×W to a first feature map of C′×H×W; A feature adjustment unit, configured to perform matrix multiplication of the first feature map and the attention map; A second convolution unit is used to perform grouped convolution and normalization on the multiplication result, and output a second feature map with a dimension of C×H×W; The residual connection unit is used to add the input feature map to the output of the second convolutional unit to form a forward-feedback feature interaction mechanism and generate a first feature map to be processed with a dimension of 2C×H×W.

3. The Transformer model for real-time image detection according to claim 1, characterized in that: The deformable attention mechanism module includes: A reference point generating unit, configured to generate N evenly distributed grid points in the first feature map to be processed as reference points, wherein the grid size is determined based on a downsampling factor; The offset prediction unit is used to predict the offset of each reference point through a lightweight sub-network and scale the offset by a predefined factor to obtain the position of the deformation point; A feature sampling unit is used to sample features at the deformation point using a bilinear interpolation function to generate keys and values; The multi-head attention unit is used to perform multi-head attention calculation on the sampled features and introduce relative position offset to form a second feature map to be processed.

4. A Transformer model for real-time image detection according to claim 3, characterized in that: In the reference point generating unit: The reference points are linearly spaced two-dimensional coordinates normalized to the range of [-1, 1]; The grid size satisfies the formula s=H / d×W / d, where d is the downsampling factor, and H and W are the feature map height and width, respectively.

5. The Transformer model for real-time image detection according to claim 3, characterized in that: In the multi-head attention unit, multi-head attention calculation is performed on the sampled features, and relative position offset is introduced, including: Set up multiple attention heads, each of which independently predicts deformation offset and attention weight; Introduce locality constraints to limit the sampling range of each reference point to a 3×3 neighborhood; Absolute position encoding and relative position deviation table are integrated in attention calculation.

6. The Transformer model for real-time image detection according to claim 1, characterized in that In the attention upsampling module: The first processing uses 1×1 transposed convolution to improve feature resolution; The second path of processing sequentially performs an upsampling operation and a 1×1 convolution operation; The third processing step sequentially performs global pooling, 1×1 convolution, and Hardsigmoid activation function processing.

7. The Transformer model for real-time image detection according to claim 1, characterized in that It also includes a hybrid aggregation network module to enhance multi-scale feature fusion by: Deploy a multi-level feature pyramid structure to integrate semantic information at different scales; Use cross-layer skip connections to reduce the feature degradation problem of deep networks; The gradient propagation efficiency of the feature pyramid is optimized through an adaptive weight distribution mechanism.

8. The Transformer model for real-time image detection according to claim 7, characterized in that: In the hybrid aggregation network module, the cross-layer jump connection is specifically: The shallow feature map is aligned with the channel dimension through 1×1 convolution and then added element-by-element to the deep feature map; The weights of the skip connections are dynamically adjusted through a learnable gating mechanism to suppress redundant feature interference; The fused feature map undergoes 3×3 convolution to further optimize the spatial information expression.

9. An application of the Transformer model for real-time image detection according to any one of claims 1 to 8 in SAR image target detection, characterized in that: When the model is applied to SAR image target detection, the following steps are included: S1. Input the SAR image to the bidirectional receptive field optimization module, extract spatial features using a k×k convolution kernel, and use a residual connection unit to suppress sea clutter interference, outputting a first feature map to be processed with a dimension of 2C×H×W; S2. Input the first feature map to be processed into the deformable attention mechanism module, generate a reference point normalized to the range of [-1, 1], focus on the scattering area of ​​the ship, aircraft or vehicle target through bilinear interpolation sampling, and output a dynamically adjusted second feature map to be processed; S3. Perform three-way processing on the dynamically adjusted second feature map to be processed by the attention upsampling module; S4. Input the target feature map processed in step S3 into the hybrid aggregation network module, fuse the multi-scale features through cross-layer skip connections, and output the final detection result, which includes the target category and bounding box coordinates.

Citation Information

Patent Citations

  • High-efficiency insulator detection system in complex space environment

    CN114399628A

  • Three-dimensional space occupation identification method and system in automatic driving scene

    CN117557985A

  • SAR image small-size ship detection method based on end-to-end Transform architecture

    CN118351453A

  • Improved target detection method in automatic driving scene based on RT-DETR

    CN118644824A

Cited By

  • Hyperspectral remote sensing image cloud detection method

    CN121121504A

  • A hyperspectral remote sensing image cloud detection method

    CN121121504B

  • Lightweight multi-scale deformable CNN-Transformer double-branch network for smoke detection

    CN121236536A