Method for detecting low-confidence small target in radar echo based on hybrid architecture
Through a radar echo detection method based on a hybrid architecture, the Hourglass3D module and the YOLOv8 network are combined to achieve multi-scale spatiotemporal feature fusion and dynamic threshold optimization, which solves the problems of accuracy and missed detection rate in small target detection in radar echoes and improves detection performance.
Patent Information
- Application Number
- CN202511344638.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing methods for detecting small targets in radar echoes have problems such as the difficulty of traditional two-dimensional convolution in capturing spatiotemporal correlations, the insufficient ability of the YOLOv8 network to extract small target features, the difficulty of fixed thresholds in the post-processing stage to adapt to changes in target density, and the lack of radar spectrum characteristic optimization, resulting in a high missed detection rate of low-confidence targets and degraded detection performance.
A radar echo detection method with a hybrid architecture generates a five-dimensional tensor through spectrum shifting and dimensionality reorganization. Combined with the Hourglass3D module and the YOLOv8 network, it realizes multi-scale spatiotemporal feature fusion, and adopts a dynamic confidence threshold and an improved non-maximum suppression algorithm to improve detection performance.
It effectively improves the detection performance of small targets in radar echoes, increases detection accuracy and reduces missed detection rate, especially in low signal-to-noise ratio environments, significantly improving the detection effect of targets.
Smart Images

Figure CN120831645A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of radar signal processing, and particularly relates to a radar echo low-confidence small target detection method based on a hybrid architecture. BACKGROUND
[0002] Small target detection in radar echo is an important research direction in target recognition field. At present, the target detection method based on deep learning (such as YOLO, Faster R-CNN, etc.) has been widely used in radar signal processing. These methods mainly extract spatial features through two-dimensional convolutional neural network, and combine post-processing algorithms such as non-maximum suppression (NMS) to optimize the detection results. In addition, for the processing of time series signals, some studies use 3D CNN or LSTM network to extract features of continuous frame radar echo to improve the detection performance of small targets. In the prior art, single-stage detectors such as YOLOv8 are preferred due to their high efficiency, but their default architecture is mainly designed for optical images, and the adaptability to low signal-to-noise ratio targets in radar echo is limited.
[0003] However, the existing methods still have significant deficiencies in radar echo small target detection: on the one hand, traditional two-dimensional convolution cannot effectively capture the spatio-temporal correlation of radar echo, resulting in a high miss detection rate of low-confidence targets; on the other hand, the backbone structure of the standard YOLOv8 network is insufficient for feature extraction of small targets, and the fixed threshold NMS in the post-processing stage cannot adapt to the dynamic changes of radar target density. In addition, the existing methods usually do not optimize for the spectral characteristics of radar data, resulting in a decline in the detection performance of small targets with low signal-to-noise ratio. SUMMARY
[0004] The purpose of the present application is to provide a radar echo low-confidence small target detection method based on a hybrid architecture to solve the problem of how to effectively improve the detection performance and accuracy of small targets with low model detection confidence caused by weak feature information in radar echo images.
[0005] The present application achieves the above-mentioned purposes through the following technical solutions: In a first aspect, the present application provides a radar echo low-confidence small target detection method based on a hybrid architecture, which is used to detect targets with confidence lower than a preset threshold. The method comprises: performing spectral shifting and dimension reorganization on the radar echo data to be measured to generate a five-dimensional tensor with time sequence characteristics; inputting the five-dimensional tensor into a detection model to obtain a prediction result including the predicted position of the target object; The detection model comprises a cascaded Hourglass3D module and a YOLOv8 network, the Hourglass3D module extracts and fuses multi-scale space-time features in the five-dimensional tensor through three-dimensional convolution and jump connection, and outputs a feature map to the front end of the backbone network of the YOLOv8 network, so as to realize multi-scale space-time feature fusion and realize classification and positioning of the target object.
[0006] Further, the dimension reorganization comprises: The storage path of the image data set obtained by symmetrically shifting the radar echo data spectrum is parsed and decoded into a NumPy array; The image data of the NumPy array is converted into RGB channel order storage; The stored data is subjected to tensor dimension rearrangement, and the channel dimension is adjusted from [H, W, C] to [C, H, W]; A batch dimension and a time dimension are added to the tensor to form a five-dimensional tensor with a [B, T, C, H, W] structure; Wherein, B is the batch size, T is the time step, C is the number of channels, H is the height, and W is the width.
[0007] Further, the Hourglass3D module comprises: An encoder composed of multiple three-dimensional convolution downsampling units, each level comprising a three-dimensional convolution layer, a normalization layer and an activation function; A bottleneck layer using three-dimensional convolution to keep the number of feature map channels unchanged; A decoder composed of multiple three-dimensional transpose convolution upsampling units, the upsampling unit comprising a 3D transpose convolution layer, a batch normalization layer and a ReLU activation function layer; A jump connection structure for inputting and fusing the output feature map of each layer of the encoder and the feature map of the corresponding layer of the decoder.
[0008] Further, the three-dimensional convolution downsampling unit uses a convolution kernel with a step size greater than 1 to realize feature map downsampling; the three-dimensional transpose convolution upsampling unit uses a convolution kernel with a step size greater than 1 to realize feature map upsampling; the jump connection structure realizes element-wise addition of the output of the nth level of the encoder and the input of the (4-n)th level of the decoder, n∈{1,2,3,4}.
[0009] Further, in the Hourglass3D module, the three-dimensional convolution downsampling unit satisfies the channel number doubling rule: the output channel number C k of the kth level is C0×2 k , where C0is the initial channel number, k∈{1,2,3,4}; the three-dimensional transpose convolution upsampling unit satisfies the channel number halving rule: the output channel number C m of the mth level is C4 / 2.(5-m) wherein C4 is the number of bottleneck layer channels, m e {1,2,3,4}.
[0010] Further, the YOLOv8 network comprises: a backbone network comprising at least two initial convolutional layers and multi-level convolutional modules, receiving the feature map output by the Hourglass3D module to the first two convolutional layers, outputting multi-scale feature maps after extracting the hierarchical representation of local features and global features by the multi-level convolutional modules; a neck network comprising a feature pyramid structure, a path aggregation structure and a dynamic convolutional layer, the feature pyramid structure fuses deep high semantic features with shallow high resolution features through an upsampling operation, the path aggregation structure realizes cross-level feature interaction through a downsampling operation, and the dynamic convolutional layer adjusts the feature fusion mode according to the adaptive weight of the input features to perform multi-scale fusion on the feature maps output by the backbone network by upsampling and downsampling; a detection head comprising a classification branch and a regression branch: the classification branch outputs the class confidence of the target object through a convolutional layer, and the regression branch outputs the prediction parameters of the target object prediction box through an independent convolutional layer.
[0011] Further, the method further comprises: setting a dynamic confidence threshold and an IoU threshold, and adaptively adjusting the threshold parameters according to the density characteristics of the target object detection scene; using an improved non-maximum suppression algorithm to process the prediction box, comprising: performing feature map interpolation enhancement on the prediction box with an area smaller than a preset threshold; outputting a detection result containing the target object class, confidence and bounding box coordinates, wherein the bounding box parameters include center point coordinates (x, y) and width and height size (w, h).
[0012] Further, the method further comprises training the cascaded Hourglass3D module and YOLOv8 network, comprising: inputting the preprocessed five-dimensional tensor [B, T, C, H, W] into the cascaded Hourglass3D module and YOLOv8 network; the Hourglass3D module processes the input data through four layers of three-dimensional convolution downsampling units, and outputs feature maps with the size halved layer by layer and the number of channels doubled layer by layer; inputting the feature map output by the last layer of the Hourglass3D module into the first two convolutional layers of the YOLOv8 network; in the backbone network of the YOLOv8 network, processing the input feature map through the multi-level convolutional module to output feature maps of three scales; inputting the feature maps of three scales into the neck network for feature fusion; The fused feature map is processed using a decoupled detection head, a classification branch outputs a category probability, and a regression branch outputs a bounding box coordinate. A classification loss and a regression loss are calculated, network parameters are updated through back propagation, and the detection model is obtained.
[0013] In a second aspect, the application provides a low-confidence small target detection system based on a hybrid architecture for radar echoes, which is used to implement the small target detection method described above, and the system comprises: A preprocessing module is configured to perform spectral shifting and dimension reorganization on the radar echo data to be measured to generate a five-dimensional tensor with time sequence characteristics. A target detection module is configured to input the five-dimensional tensor into a detection model to obtain a prediction result including a predicted bounding box position of a target object. The detection model comprises a cascaded Hourglass3D module and a YOLOv8 network, the Hourglass3D module extracts and fuses multi-scale spatio-temporal features in the five-dimensional tensor through three-dimensional convolution and skip connection, and outputs a feature map to the front end of the backbone network of the YOLOv8 network to realize multi-scale spatio-temporal feature fusion and classification and positioning of the target object.
[0014] Further, the system further comprises: An optimization processing module is configured to set a dynamic confidence threshold and an IoU threshold, and adaptively adjust the threshold parameters according to the density characteristics of the target object detection scene; an improved non-maximum suppression algorithm is used to process the predicted bounding box, including: performing feature map interpolation enhancement on the predicted bounding box with an area smaller than a preset threshold; and outputting a detection result containing a target object category, a confidence, and a bounding box coordinate, wherein the bounding box parameter comprises a center point coordinate (x, y) and a width and height size (w, h).
[0015] The application has the following advantages: 1. The application constructs a hybrid architecture by cascading the Hourglass3D module and the YOLOv8 network, effectively improving the detection performance of small targets in radar echoes. The Hourglass3D module uses three-dimensional convolution operation to process time sequence features, realizes multi-scale spatio-temporal feature extraction and fusion through an encoder-bottleneck layer-decoder structure and skip connection, and enhances the feature expression ability of small targets. The multi-scale feature extraction and FPN+PAN feature fusion mechanism of the YOLOv8 network further optimizes the detection effect of targets of different scales.
[0016] 2、The application preserves the space-time feature information of radar echo by the data organization form of five-dimensional tensor [B, T, C, H, W], so that the network can fully utilize the correlation between continuous frames. The decoupling detection head design realizes independent optimization of classification and positioning tasks, and improves the accuracy of small target detection. The dynamic non-maximum suppression strategy and small target compensation mechanism effectively reduce the missed detection and false detection. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 A flowchart of the low-confidence small target detection method in the radar echo based on the mixed architecture provided by the embodiment of the application is provided. Figure 2 Another flowchart of the low-confidence small target detection method in the radar echo based on the mixed architecture provided by the embodiment of the application is provided. Figure 3 A structural schematic diagram of the Hourglass3D module provided by the embodiment of the application is provided. Figure 4 A comparison chart of the Precision-Recall curves of the YOLOv8 model and the HG3DCN-YOLOv8 model in the simulation analysis part of the application is provided. Figure 5 A comparison chart of the Precision-Confidence curves of the YOLOv8 model and the HG3DCN-YOLOv8 model in the simulation analysis part of the application is provided. Figure 6 A schematic diagram of manually annotated original label in the data set in the simulation analysis part of the application is provided. Figure 7 A schematic diagram of the detection result of the target detection of the YOLOv8 model in the simulation analysis part of the application is provided. Figure 8 A schematic diagram of the detection result of the target detection of the HG3DCN-YOLOv8 model in the simulation analysis part of the application is provided. DETAILED DESCRIPTION
[0018] The following further describes the application in conjunction with the drawings. It is necessary to point out here that the following detailed description is only used to further illustrate the application, and cannot be understood as limiting the protection scope of the application. Those skilled in the art can make some non-essential improvements and adjustments to the application according to the above application content.
[0019] It's worth noting that low-confidence small targets in radar echoes are those difficult to accurately detect using traditional methods due to their small reflection cross-section and low signal-to-noise ratio, such as drones or stealth targets. Existing methods based on two-dimensional convolutional neural networks (such as YOLOv8) have significant drawbacks: First, they struggle to effectively capture the spatiotemporal correlations of radar echoes, leading to missed detection of dynamic targets; second, the network's downsampling process easily misses the subtle features of small targets; third, fixed-threshold post-processing algorithms cannot adapt to dynamic changes in target density; and finally, the lack of optimized design for radar spectral characteristics further degrades detection performance.
[0020] Example 1 like Figures 1-3 As shown, in response to the above-mentioned defects, a specific embodiment of the present application proposes a low-confidence small target detection method in radar echoes based on a hybrid architecture, which is used to detect targets with confidence lower than a preset threshold. The method includes: performing spectral shifting and dimensional reorganization on the radar echo data to be measured to generate a five-dimensional tensor with time series characteristics; inputting the five-dimensional tensor into the detection model to obtain a prediction result including the position of the target prediction box; wherein, the detection model includes a cascaded Hourglass3D module and a YOLOv8 network, and the Hourglass3D module extracts and fuses multi-scale spatiotemporal features in the five-dimensional tensor through three-dimensional convolution and jump connection, and outputs a feature map to the backbone network front end of the YOLOv8 network to realize multi-scale spatiotemporal feature fusion and classification and positioning of the target.
[0021] In this application, please combine Figure 2 The Hourglass3D module is an improved design based on the classic Hourglass network architecture, used to extract the joint spatiotemporal features of radar echoes. This module adopts an encoder-bottleneck layer-decoder structure. The encoder gradually compresses the spatial dimension and expands the number of channels through multi-level 3D convolution downsampling. The bottleneck layer maintains the number of feature map channels unchanged through 3D convolution to capture local contextual information. The decoder restores the spatial resolution through 3D transposed convolution upsampling and adds the feature maps of the corresponding encoder layer element-by-element through skip connections to achieve multi-scale feature fusion.
[0022] Specifically, the dimension reorganization includes: the radar echo data spectrum can be symmetrically moved to obtain an image data set by using the fftshift tool in MATLAB, the storage path of the image data set is parsed, and is decoded into a NumPy array, which is a structured numerical container generated by the Python scientific computing library NumPy, and is used for storing the preprocessed radar echo data; the image data of the NumPy array is converted to be stored in the RGB channel order; tensor dimension rearrangement is performed on the stored data, and the channel dimension [C] is moved from the last dimension [H, W, C] to the first dimension [C, H, W]; then, two additional dimensions, a batch dimension (Batch) and a time dimension (Time), are added; finally, the tensor shape changes from [C, H, W] to [B, T, C, H, W], which respectively correspond to the batch size (the number of samples processed in parallel during training), the time step (the number of consecutive radar frames), the number of channels (corresponding to the radar frequency band / polarization information), the height, and the width; H / W is the spatial resolution.
[0023] It can be understood that the five-dimensional tensor [B, T, C, H, W] generated by the application completely retains the space-time characteristic information of the radar echo by integrating the batch, time sequence, channel, and spatial dimensions: the batch dimension (B) supports parallel training, the time dimension (T) captures the target motion trajectory in the continuous frame sequence, the channel dimension (C) stores the multi-frequency / polarization radar features, and the spatial dimension (H, W) maintains the geometric structure of the target; as the standardized input of the Hourglass3D module, the tensor directly processes the time-space correlation features through a three-dimensional convolution kernel, enabling the network to simultaneously analyze the instantaneous scattering characteristics and cross-frame evolution rules of the target, especially for small targets with a signal-to-noise ratio less than 10 dB and a pixel area less than 0.5%.
[0024] Further, the Hourglass3D module includes an encoder, a bottleneck layer, a decoder, and a skip connection structure; the encoder is composed of multiple three-dimensional convolution downsampling units, each level containing a three-dimensional convolution layer, a normalization layer, and an activation function; the bottleneck layer uses three-dimensional convolution to keep the number of feature map channels unchanged; the decoder is composed of multiple three-dimensional transpose convolution upsampling units, and the upsampling unit contains a 3D transpose convolution layer, a batch normalization layer, and a ReLU activation function layer; the skip connection structure fuses the output feature maps of each layer of the encoder with the input feature maps of the corresponding level of the decoder.
[0025] Among them, the three-dimensional convolution downsampling unit uses a convolution kernel with a step size greater than 1 to realize feature map downsampling (corresponding to downsampling); the three-dimensional transpose convolution upsampling unit uses a convolution kernel with a step size greater than 1 to realize feature map upsampling (corresponding to upsampling); the skip connection structure adds the output of the nth level of the encoder to the input of the (4-n)th level of the decoder element by element, n e {1, 2, 3, 4}.
[0026] As a preferred solution, in the Hourglass3D module, the three-dimensional convolution downsampling unit satisfies the channel number multiplication rule: the output channel number C k = C0×2 k , where C0 is the initial channel number, k ∈ {1, 2, 3, 4}; the three-dimensional transpose convolution upsampling unit satisfies the channel number reduction rule: the output channel number C m = C4 / 2 (5-m) , where C4 is the bottleneck layer channel number (i.e. the last stage of the encoder C4 = C0×2 4 ), m ∈ {1, 2, 3, 4}. Wherein, k ∈ {1, 2, 3, 4}, corresponding to the 4-stage downsampling unit of the encoder (i.e. k = 1 is the first stage of downsampling, k = 4 is the fourth stage of downsampling); similarly, m corresponds to the 4-stage upsampling unit of the decoder.
[0027] In this application, the Hourglass3D module realizes multi-scale spatio-temporal feature extraction through the cascaded structure of three-dimensional convolution downsampling unit and three-dimensional transpose convolution upsampling unit.
[0028] Specifically, the three-dimensional convolution downsampling unit uses a 3x3x3 convolution kernel with a stride greater than 1 to downsample the feature map. After each stage of downsampling, the spatial size (H, W) of the feature map is halved, and the channel number is multiplied according to the rule C k = C0×2 k (k ∈ {1, 2, 3, 4}), for example, when the input feature map size is [8, 5, 64, 256, 256], after the first stage of downsampling, it becomes [8, 5, 128, 128, 128], this design compresses the spatial information while enhancing the expression ability of the feature. The corresponding three-dimensional transpose convolution upsampling unit uses a 3x3x3 transpose convolution kernel with a stride greater than 1 to realize feature map upsampling, so that the spatial size is recovered to the original resolution level by level, and the channel number is reduced according to the rule C m = C4 / 2 (5-m) (m ∈ {1, 2, 3, 4}), for example, the [8, 5, 512, 16, 16] feature map output by the bottleneck layer becomes [8, 5, 256, 32, 32] after the first stage of upsampling.
[0029] In particular, the output of the nth stage of the encoder is element-wise added to the input of the (4-n)th stage of the decoder through a skip connection structure (n e {1, 2, 3, 4}), for example, the [8, 5, 128, 128, 128] feature map of the output of the first stage of the encoder is added to the [8, 5, 128, 128, 128] feature map of the output of the third stage of the decoder, realizing the complementary advantages of low-level detail features and high-level semantic features. This symmetrical encoding and decoding structure cooperates with the cross-layer feature fusion mechanism to effectively solve the feature disappearance problem caused by downsampling of small targets in radar echoes.
[0030] As a preferred solution, the YOLOv8 network comprises: The backbone network comprises at least two initial convolutional layers and multi-stage convolutional modules, receives the feature map output by the Hourglass3D module to the first two convolutional layers, and outputs a multi-scale feature map after extracting the hierarchical representation of local features and global features through the multi-stage convolutional modules.
[0031] The neck network comprises a feature pyramid structure, a path aggregation structure, and a dynamic convolutional layer. The feature pyramid structure fuses deep high semantic features and shallow high resolution features through upsampling operations. The path aggregation structure realizes cross-level feature interaction through downsampling operations. The dynamic convolutional layer adjusts the feature fusion mode according to the adaptive weight of the input features, and performs multi-scale fusion on the feature map output by the backbone network using upsampling and downsampling.
[0032] The detection head comprises a classification branch and a regression branch. The classification branch outputs the class confidence of the target object through a convolutional layer. The regression branch outputs the prediction parameters of the target object prediction box through an independent convolutional layer.
[0033] In this application, the YOLOv8 network as a detection core module realizes efficient detection of small radar targets through the following optimization design: After the backbone network receives the five-dimensional feature map output by the Hourglass3D module, it first performs channel dimension adaptation and shallow feature extraction through two initial convolutional layers (kernel size 3x3, step 1), and the output feature map size remains consistent with the input. Subsequently, deep feature extraction is performed through a residual structure composed of multi-stage CBS modules (Conv-BN-SiLU). Each stage uses convolution kernels of different strides (1x1 and 3x3 combination) to realize hierarchical capture of local details and global semantics, and finally outputs three scales of feature maps (for example, when the input is 640x640, the output is 80x80, 40x40, 20x20 three resolutions), corresponding to different sizes of target detection requirements.
[0034] In particular, the first convolutional module preserves the spatiotemporal feature dimensions of the output of the Hourglass3D, and the subsequent modules gradually compress the temporal information and strengthen the spatial features.
[0035] The neck network adopts an improved FPN+PAN dual-channel structure: the FPN channel adds and fuses deep high semantic features (such as 20x20) and shallow high-resolution features (such as 40x40) element by element through 2 times up-sampling; the PAN channel realizes reverse feature enhancement through 3x3 convolution down-sampling with a step of 2. The dynamic convolution layer automatically adjusts the convolution kernel weight according to the energy distribution of the feature map, for example, a larger receptive field convolution kernel (5x5 instead of 3x3) is used for small target object dense areas (determined by threshold segmentation).
[0036] The detection head adopts a decoupling design: the classification branch outputs the category confidence through 1x1 convolution, and the regression branch predicts the bounding box parameters (the center coordinates x and y are normalized by sigmoid, and the width and height w and h are predicted based on the logarithmic offset of the predicted box) through 3x3 deep separable convolution, and the positioning accuracy is optimized through CIoU Loss.
[0037] As a preferred solution, the method further comprises: setting a dynamic confidence threshold and an IoU threshold, and adaptively adjusting the threshold parameters according to the density characteristics of the target object detection scene; using an improved non-maximum suppression algorithm to process the predicted box, including: performing feature map interpolation enhancement on the predicted box with an area smaller than a preset threshold; outputting the detection result containing the target object category, confidence and bounding box coordinates, wherein the bounding box parameters include the center point coordinates (x, y) and the width and height size (w, h).
[0038] The present application adopts a dynamic optimization strategy in the post-processing stage to improve the performance of small target object detection. In specific implementation, the dynamic confidence threshold is automatically adjusted according to the spatial density distribution of the target object: by counting the clustering of the predicted box in the current frame, for example, the confidence threshold is increased from the baseline value of 0.4 to 0.5-0.6 for dense areas (such as target objects with a distance of less than 10 pixels), and decreased to 0.3-0.35 for sparse areas, effectively balancing the missed detection and false detection. The improved non-maximum suppression algorithm is specially optimized for small target objects: first, for the predicted box with an area smaller than 32x32 pixels, perform bilinear interpolation enhancement (weight fusion of 4x4 neighborhood feature values) at its corresponding feature map position, which greatly improves the feature response of weak target objects; then, a dynamic IoU threshold strategy is used to adaptively adjust the suppression threshold according to the target object density. The bounding box parameter output uses normalized coordinates.
[0039] As a preferred solution, the method further comprises training the cascaded Hourglass3D module and the YOLOv8 network, including: inputting the preprocessed five-dimensional tensor [B, T, C, H, W] into the cascaded Hourglass3D module and the YOLOv8 network; the Hourglass3D module processes the input data through four three-dimensional convolution downsampling units, and outputs feature maps whose sizes are halved layer by layer and whose channel numbers are multiplied layer by layer; the feature map output by the last layer of the Hourglass3D module is input into the first two convolution layers of the YOLOv8 network; in the backbone network of the YOLOv8 network, the input feature map is processed through a multi-level convolution module, and three scale feature maps are output; the three scale feature maps are input into the neck network for feature fusion; the fused feature map is processed using a decoupled detection head, the classification branch outputs class probability, and the regression branch outputs bounding box coordinates; the classification loss and the regression loss are calculated, the network parameters are updated through back propagation, and a detection model is obtained.
[0040] In specific implementation, the joint training process of the cascaded Hourglass3D module and the YOLOv8 network is implemented through the following specific steps: During training, the three-dimensional radar echo sequence (five-dimensional tensor [B, T, C, H, W]) generated by preprocessing is input into the network. The Hourglass3D module first processes through four three-dimensional convolution downsampling, each layer uses a 3×3×3 convolution kernel (step size 2, zero padding 1), and cooperates with the channel number multiplication rule to compress the feature map size from [8, 5, 64, 256, 256] to [8, 5, 1024, 16, 16] level by level, while retaining the feature maps of each level through a jump connection. The spatio-temporal features output by the module are adjusted in channel number through a 1×1×1 convolution, and then input into the first two convolution layers (kernel size 3×3, step size 1) of the YOLOv8 network, which is connected with the original backbone network of the YOLOv8. The backbone network processes through four CBS modules (Conv-BN-SiLU), and outputs feature maps of three scales of 20×20, 40×40 and 80×80, which correspond to target detection tasks of different sizes respectively. The neck network adopts an improved FPN+PAN structure, wherein the FPN path fuses deep semantic features through bilinear interpolation upsampling, the PAN path enhances positioning accuracy through 3×3 convolution (step size 2) downsampling, and the dynamic convolution layer automatically adjusts the fusion weight according to the feature energy distribution. In the decoupled detection head, the classification branch uses a 1×1 convolution to output class probability, and uses Focal Loss to solve the sample imbalance; the regression branch predicts the bounding box through a 3×3 depth separable convolution, and uses CIoU Loss to optimize positioning, wherein the center coordinates (x, y) are normalized through sigmoid, and the width and height (w, h) are calculated based on the logarithmic offset of the anchor box size.
[0041] Specifically, during model training, the data is divided into training and validation sets in a ratio of 8:2, the batch size is set to 8, the number of iterations is 100 rounds, and the early stopping mechanism is set, using the patience parameter. If there is no improvement in the monitoring indicator mAP@0.5:0.95 in the next thirty rounds of training, stop training to avoid overfitting. After each round of training, the performance of the validation set (HG3D is the abbreviation of Hourglass3D, and CN is the abbreviation of Cascade Network level networking) is evaluated, and the loss and accuracy on the validation set are monitored.
[0042] In addition, the average precision mean mAP, F1 score, precision P and recall R indicators are calculated, the Precision-Confidence curve and Precision-Recall curve are drawn, and the dataset samples and detection results are visualized to determine their performance in real situations. Among them, the average precision mean is an important indicator for measuring the detection accuracy of multi-class targets; the F1 score is the weighted average of precision and recall, which is used to measure the comprehensive performance and stability of the model; the precision is used to measure the proportion of true positive examples in the samples predicted as positive examples by the model, reflecting the accuracy of the model prediction; the recall is used to measure the proportion of correctly predicted positive examples in all actual positive examples, reflecting the capture ability of the model to positive examples.
[0043] According to the above embodiment, the working principle of the application is as follows: The application proposes the core idea of joint extraction and multi-scale fusion based on space-time features, which completely preserves the time sequence dynamic characteristics and spatial structure information of radar echoes through spectral symmetric shift and five-dimensional tensor reorganization. The Hourglass3D module adopts a symmetric encoding and decoding structure. The encoder compresses the spatial dimension and expands the channel number step by step through three-dimensional convolution downsampling, while preserving the target motion trajectory and enhancing the feature expression ability. The decoder restores the resolution through transposed convolution and combines with the jump connection to fuse multi-scale features, effectively solving the feature loss problem of small targets in the downsampling process. The YOLOv8 network extracts multi-level features through the backbone network, the neck network adopts a bidirectional feature pyramid to realize deep and shallow feature complementation, and the dynamic convolution layer adaptively adjusts the receptive field according to the target distribution. The decoupling design of the detection head makes the classification and positioning tasks promote each other, and the dynamic threshold strategy optimizes the post-processing process. The whole system realizes the cooperative optimization of space-time features through end-to-end training, so that the network can capture the instantaneous scattering characteristics and cross-frame evolution law of the target at the same time, significantly improving the detection performance of small targets in low signal-to-noise ratio environments.
[0044] Example 2 Based on the same inventive concept, a specific embodiment of the present application proposes a radar echo low-confidence small target detection system based on a hybrid architecture, which is used to implement the small target detection method proposed in Embodiment 1. The system includes a preprocessing module, a target detection module, and an optimization processing module. The preprocessing module is used to perform spectral shifting and dimension reconstruction on the radar echo data to be measured to generate a five-dimensional tensor with time sequence characteristics. The target detection module is used to input the five-dimensional tensor into a detection model to obtain a prediction result including the predicted position of the target object. The detection model includes a cascaded Hourglass3D module and a YOLOv8 network. The Hourglass3D module extracts and fuses multi-scale space-time features in the five-dimensional tensor through three-dimensional convolution and skip connection, and outputs the feature map to the front end of the backbone network of the YOLOv8 network to realize multi-scale space-time feature fusion and target object classification and positioning.
[0045] The optimization processing module is used to set dynamic confidence threshold and IoU threshold, and adaptively adjust the threshold parameters according to the density characteristics of the target detection scene. The improved non-maximum suppression algorithm is used to process the prediction box, including: performing feature map interpolation enhancement on the prediction box with an area smaller than a preset threshold; and outputting a detection result containing the target category, confidence, and bounding box coordinates, wherein the bounding box parameters include the center point coordinates (x, y) and the width and height size (w, h).
[0046] For specific limitations of the radar echo low-confidence small target detection system based on the hybrid architecture, refer to the limitations of the radar echo low-confidence small target detection method based on the hybrid architecture in the foregoing, which will not be repeated here. It should be noted that each module in the above detection system corresponds to each step in the implementation of the above detection method, and the instances and application scenarios realized by the multiple modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1.
[0047] It can be understood that the radar echo low-confidence small target detection system proposed by the present application adopts a three-level processing architecture: the preprocessing module is responsible for spectral shifting and five-dimensional tensor conversion of radar data, and reorganizes the original echo data into a [B, T, C, H, W] structure containing time sequence characteristics; the target detection module realizes feature extraction and recognition through a cascaded Hourglass3D and YOLOv8 network, wherein the Hourglass3D module uses three-dimensional convolution to mine space-time correlation features, and the YOLOv8 network completes multi-scale target detection; the optimization processing module adopts a dynamic threshold strategy to automatically adjust the confidence and IoU threshold according to the target object density, and performs feature enhancement on the small target object prediction box.
[0048] In specific applications, the above detection system is suitable for real-time detection scenes such as border monitoring and unmanned aerial vehicle early warning.
[0049] In order to make the present application and its advantages clearer, the following will further explain the method provided by the present application in combination with specific simulation experiments and related partial diagrams.
[0050] (1) The training time, mAP@0.5 and F1 score index results of the YOLOv8 model and the HG3DCN-YOLOv8 model in the present application are shown in Table 1.
[0051] Table 1 Training time, mAP@0.5 and F1 score index results ; The time used for training the improved model in the present application is more than ten times less than the training time of the original YOLOv8 model, and the calculation efficiency is greatly improved. At the same time, the mAP@0.5 is also increased by 2.89%, indicating that the improved model has improved the accuracy of target positioning and reduced the missed detection rate of small targets.
[0052] (2) Figure 4 The Precision-Recall curves of the YOLOv8 model and the HG3DCN-YOLOv8 model are compared, where the horizontal coordinate represents the recall rate and the vertical coordinate represents the precision. It can be seen from the figure that the precision of the HG3DCN-YOLOv8 model is higher than that of the original YOLOv8 model, especially when the recall rate R is small, that is, in the case where the target detection is more difficult, the improved model in the present application detects the target more accurately, that is, has better reliability.
[0053] The skip connection mechanism of Hourglass3D in the present application preserves the low-level features of small targets, and the dynamic FPN-PAN structure optimizes the multi-scale feature fusion.
[0054] (3) Figure 5 The Precision-Confidence curves of the YOLOv8 model and the HG3DCN-YOLOv8 model are compared, where the horizontal coordinate represents the confidence and the vertical coordinate represents the precision. The curves show that the HG3DCN-YOLOv8 model not only has improved precision at a confidence of 0.5, but also has better performance in the entire confidence interval from 0 to 1.
[0055] (4) Figure 6 , Figure 7 and Figure 8 are the original labels manually labeled in the data set, the detection results of target detection using YOLOv8, and the detection results of target detection using the improved model, respectively, Figure 7 It can be seen that YOLOv8 misdetected a single small target as two adjacent targets, which is due to the difficulty of two-dimensional convolution in distinguishing the spatiotemporal features of dense targets.Figure 8 The middle HG3DCN-YOLOv8 correctly detects the target object because the three-dimensional convolution captures the consistency of the target object motion between consecutive frames, the dynamic NMS suppresses redundant frames, and the feature interpolation enhances the response signal of the small target object.
[0056] In summary, the improved Hourglass3D module is used to capture the motion features between consecutive frames, and the dynamic optimized YOLOv8 detection network is used to realize multi-scale target object recognition. The detection reliability of small target objects is significantly improved through feature enhancement and adaptive threshold strategy. Simulation verification shows that compared with the traditional method, the scheme has obvious advantages in detection accuracy, training efficiency and robustness, and is especially suitable for complex radar monitoring scenes with low signal-to-noise ratio and small target size.
[0057] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0058] In addition, the functional modules in each of the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0059] The above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting low-confidence small targets in radar echoes based on a hybrid architecture, the method being used for detecting target objects with a confidence lower than a preset threshold, and characterized in that, The method comprises: Spectrum shifting and dimension reorganization are performed on the radar echo data to be tested to generate a five-dimensional tensor with timing characteristics; The five-dimensional tensor is input into a detection model to obtain a prediction result including a target object prediction box position; The detection model comprises a cascaded Hourglass3D module and a YOLOv8 network, the Hourglass3D module extracts and fuses multi-scale space-time features in the five-dimensional tensor through three-dimensional convolution and jump connection, and outputs a feature map to the front end of the backbone network of the YOLOv8 network to realize multi-scale space-time feature fusion and target object classification and positioning.
2. The method of claim 1, wherein, The dimension reorganization comprises: The storage path of the image data set obtained by symmetric spectrum shifting of the radar echo data is parsed and decoded into a NumPy array; The image data of the NumPy array is converted to be stored in RGB channel order; The stored data is subjected to tensor dimension rearrangement to adjust the channel dimension from [H, W, C] to [C, H, W]; Batch dimension and time dimension are added to the tensor to form a five-dimensional tensor with a [B, T, C, H, W] structure; B is the batch size, T is the time step, C is the number of channels, H is the height, and W is the width.
3. The method of claim 1, wherein, The Hourglass3D module comprises: An encoder composed of multiple three-dimensional convolution downsampling units, each level comprising a three-dimensional convolution layer, a normalization layer and an activation function; A bottleneck layer using three-dimensional convolution to keep the number of feature map channels unchanged; A decoder composed of multiple three-dimensional transpose convolution upsampling units, the upsampling unit comprising a 3D transpose convolution layer, a batch normalization layer and a ReLU activation function layer; A jump connection structure that inputs and fuses the output feature maps of each layer of the encoder with the feature maps of the corresponding level of the decoder.
4. The method of claim 3, wherein, The three-dimensional convolution downsampling unit uses a convolution kernel with a step size greater than 1 to realize feature map downsampling; the three-dimensional transpose convolution upsampling unit uses a convolution kernel with a step size greater than 1 to realize feature map upsampling; the jump connection structure adds the output of the nth level of the encoder to the input of the (4-n)th level of the decoder element by element, n∈{1,2,3,4}.
5. The method of claim 4, wherein, In the Hourglass3D module, the three-dimensional convolution downsampling unit satisfies the channel number multiplication rule: the output channel number C k of the kth stage C0×2 k , wherein C0 is the initial channel number, and k∈{1,2,3,4}; the three-dimensional transpose convolution upsampling unit satisfies the channel number division rule: the output channel number C m of the mth stage C4 / 2 (5-m) , wherein C4 is the bottleneck layer channel number, and m∈{1,2,3,4}.
6. The method of claim 1, wherein, The YOLOv8 network comprises: A backbone network comprising at least two initial convolution layers and multiple convolution modules, receiving the feature map output by the Hourglass3D module to the first two convolution layers, outputting a multi-scale feature map after extracting hierarchical representations of local features and global features through the multiple convolution modules; A neck network comprising a feature pyramid structure, a path aggregation structure and a dynamic convolution layer, the feature pyramid structure fuses deep high semantic features with shallow high resolution features through upsampling operations, the path aggregation structure realizes cross-level feature interaction through downsampling operations, and the dynamic convolution layer adjusts the feature fusion mode according to the adaptive weight of the input feature to perform multi-scale fusion on the feature map output by the backbone network through upsampling and downsampling; A detection head comprising a classification branch and a regression branch: the classification branch outputs the class confidence of the target object through a convolution layer, and the regression branch outputs the prediction parameters of the target object prediction box through an independent convolution layer.
7. The method of claim 1, wherein, The method further comprises: A dynamic confidence threshold and an IoU threshold are set, and the threshold parameters are adaptively adjusted according to the density characteristics of the target detection scene; The improved non-maximum suppression algorithm is used to process the prediction box, including: performing feature map interpolation enhancement on the prediction box with an area less than a preset threshold; The detection result containing the target class, confidence and bounding box coordinates is output, wherein the bounding box parameters include the center point coordinates (x, y) and the width and height size (w, h).
8. The method of claim 1, wherein, The method further comprises training the cascaded Hourglass3D module and YOLOv8 network, including: The preprocessed five-dimensional tensor [B, T, C, H, W] is input into the cascaded Hourglass3D module and YOLOv8 network; The Hourglass3D module processes the input data through four three-dimensional convolution downsampling units, and the feature map size is halved and the channel number is doubled layer by layer; The feature map output by the last layer of the Hourglass3D module is input into the first two convolution layers of the YOLOv8 network; In the backbone network of the YOLOv8 network, the input feature map is processed through a multi-level convolution module, and three scale feature maps are output; The three scale feature maps are input into the neck network for feature fusion; The decoupled detection head is used to process the fused feature map, and the classification branch outputs the class probability and the regression branch outputs the bounding box coordinates; The classification loss and the regression loss are calculated, and the network parameters are updated through back propagation to obtain the detection model.
9. A low confidence small target detection system in radar echo based on hybrid architecture, characterized in that, The system is used to implement the small target detection method of any one of claims 1-8, and the system comprises: A preprocessing module is configured to perform spectral shifting and dimension reconstruction on the to-be-detected radar echo data to generate a five-dimensional tensor with time sequence characteristics; A target detection module is configured to input the five-dimensional tensor into a detection model to obtain a prediction result including a target prediction box position; The detection model comprises a cascaded Hourglass3D module and a YOLOv8 network, the Hourglass3D module extracts and fuses multi-scale spatio-temporal features in the five-dimensional tensor through three-dimensional convolution and jump connection, and outputs a feature map to the front end of the backbone network of the YOLOv8 network to realize multi-scale spatio-temporal feature fusion and target classification and positioning.
10. The hybrid architecture based low confidence small target detection system in radar returns according to claim 9, wherein, The system further comprises: An optimization processing module is configured to set a dynamic confidence threshold and an IoU threshold, and adaptively adjust the threshold parameters according to the density characteristics of the target detection scene; an improved non-maximum suppression algorithm is used to process the prediction box, including: performing feature map interpolation enhancement on the prediction box with an area less than a preset threshold; and outputting a detection result containing the target class, confidence and bounding box coordinates, wherein the bounding box parameters include the center point coordinates (x, y) and the width and height size (w, h).
Citation Information
Patent Citations
Radar target intelligent detection method based on RAD domain data three-dimensional joint correlation features
CN117269924A
Radar weak target detection method and device based on deep learning
CN119575332A
Real-time Aerial Suspicious Analysis (ASANA) System and Method for Identification of Suspicious individuals in public areas
US20200394384A1
Cited By
Tablet flaw detection and elimination system and method, medium, product and terminal
CN122098950A
Tablet defect detection and rejection system, method, medium, product, and terminal
CN122098950B