A method, apparatus, equipment and medium for detecting small targets at low altitudes.
By employing cross-scale adaptive feature enhancement and modal conditional modulation processing, the semantic conflict and feature loss problems in low-altitude small target detection are resolved, achieving high-precision and real-time low-altitude small target detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TOWER CO LTD
- Filing Date
- 2026-06-05
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the detection of small targets at low altitudes suffers from semantic conflicts caused by the fusion of heterogeneous modal features and the easy loss of small target features, resulting in insufficient detection accuracy and an inability to meet the high-precision requirements of complex low-altitude scenarios.
A cross-scale adaptive feature enhancement module and modality conditional modulation processing are adopted. Multi-scale feature maps are extracted through the backbone network, semantic refinement and modality-aware modulation are performed, and feature aggregation and target detection are performed by combining a hybrid encoder and decoder.
It significantly improves the feature retention rate and detection accuracy of small targets in complex low-altitude environments, resolves semantic conflicts caused by heterogeneous modal feature fusion, enhances the system's perception robustness under fragmented data, and balances high accuracy and real-time performance.
Smart Images

Figure CN122493033A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer vision and low-altitude safety protection technology, and specifically relates to a method, device, equipment and medium for detecting small targets at low altitudes. Background Technology
[0002] Object detection is one of the core tasks in the field of computer vision, with the aim of accurately identifying and locating target objects of interest from images or videos.
[0003] In existing technologies, the structure and main working steps of visual detection for low-altitude dual-modal small targets are as follows: The backbone network performs hierarchical convolution on the input infrared or visible light image to extract multi-scale feature maps with decreasing spatial resolution and increasing semantic information; the hybrid encoder stage performs a unified linear projection on the above features and uses standard FPN for top-down feature aggregation to enhance feature flow between different levels; for heterogeneous infrared and visible light data, pixel-level concatenation or feature-level weighted summation is used for preliminary integration of modal information; simple modal cues or modulation operators are introduced to map discrete modal IDs into continuous vectors and apply conditional influences to the feature sequence before entering the decoder; the Transformer decoder receives the encoded and modulated feature sequence, interacts with the preset target query through a cross-attention mechanism, and finally outputs the target's category confidence and bounding box coordinates.
[0004] However, when facing complex real-world scenarios involving small, weak features and strong interference at low altitudes, this existing technology still has the following technical problems: 1. Significant semantic conflicts arise from the heterogeneous features of dual-modal imaging, limiting detection accuracy. Infrared and visible light images are generated based on different imaging mechanisms, resulting in inherent heterogeneity in their underlying pixel distribution, texture features, and semantic representations, with vastly different feature domains. Existing technologies only employ fusion methods such as hard weight sharing, pixel-level cascading, or simple feature weighted summation, without differentiated adaptation and calibration of heterogeneous features. This easily leads to semantic conflicts and feature distribution disturbances after the dual-modal features are superimposed. Effective complementary features are masked by redundant interference information, significantly reducing the quality of dual-modal information fusion. Ultimately, this makes it difficult to improve model detection accuracy and adapt to the high-precision detection requirements of complex low-altitude scenes.
[0005] 2. Features of tiny targets are easily lost during multi-scale downsampling. Low-altitude tiny targets have extremely low pixel proportions, sparse effective features, and weak semantic representations, requiring extremely high preservation of feature details. Existing technologies employ cross-layer downsampling and multi-scale feature interaction mechanisms in standard detection frameworks. During multi-layer convolution and downsampling dimensionality reduction, high-frequency details, edge contours, and weak local features of tiny targets are irreversibly lost. This causes the model to fail to capture key semantic clues of tiny targets, easily leading to missed detections and false detections, resulting in extremely poor performance in tiny target detection.
[0006] Therefore, there is an urgent need for a technical solution that can solve the problems of fusion of heterogeneous infrared and visible light data and the easy loss of fine features of small targets, and improve the all-weather comprehensive perception capability in complex low-altitude environments. Summary of the Invention
[0007] This application provides a method, apparatus, device, and medium for detecting small targets at low altitudes, which can solve the technical problems in the prior art, such as semantic conflicts caused by the awkward fusion of heterogeneous modal features and the inability to capture local features of small targets.
[0008] To achieve the above objectives, this application provides the following technical solution: A method for detecting small targets at low altitudes, the method comprising: The backbone network is used to extract features from the acquired image to be detected, resulting in N feature maps of different scales, where N is a positive integer. Semantic refinement is performed on the N feature maps to obtain a refined multi-scale feature map sequence. Based on the refined multi-scale feature map sequence, multi-scale feature aggregation is performed through a hybrid encoder to obtain an aggregated multi-scale feature map sequence. Based on the aggregated multi-scale feature sequence and the discrete modal identifier corresponding to the image to be detected, modal conditional modulation processing is performed to obtain modal-aware refined features. The modality-aware refined features are input into the decoder to obtain the class score and bounding box position coordinates of the low-altitude small target output by the decoder.
[0009] Based on the same inventive concept, this application also provides a low-altitude small target detection device, the device comprising: The feature extraction module is configured to extract features from the acquired image to be detected through the backbone network, and obtain N feature maps of different scales, where N is a positive integer; The feature enhancement module is configured to perform semantic refinement on the N feature maps to obtain a refined multi-scale feature map sequence. The hybrid encoding module is configured to perform multi-scale feature aggregation based on the refined multi-scale feature map sequence using a hybrid encoder to obtain an aggregated multi-scale feature map sequence. The prompt modulation module is configured to perform modality conditional modulation processing based on the aggregated multi-scale feature sequence and the discrete modality identifier corresponding to the image to be detected, to obtain modality-aware refined features; The feature decoding module is configured to input the modality-aware refined features into the decoder to obtain the class score and bounding box position coordinates of the low-altitude small target output by the decoder.
[0010] Based on the same inventive concept, this application also provides an electronic device, including: a memory and a processor; the processor is used to read and execute a computer program stored in the memory to implement the steps of the aforementioned low-altitude small target detection method.
[0011] Based on the same inventive concept, this application also provides a computer storage medium storing computer-executable instructions, which, when executed, implement the steps of the aforementioned low-altitude small target detection method.
[0012] Compared with the prior art, this application has the following advantages: On the one hand, it actively performs semantic refinement, which effectively overcomes the problem that traditional networks are prone to losing weak high-frequency features during downsampling, and significantly improves the feature retention rate and final detection accuracy of small targets (such as long-distance drones and birds) in complex low-altitude environments.
[0013] On the other hand, modal conditional modulation processing is performed to achieve adaptive perception and feature decoupling of the input modality, effectively eliminating semantic conflicts and domain shifts caused by the fusion of heterogeneous infrared and visible light data. Combined with the overall minimalist and lightweight design, it not only solves the feature oscillation problem in multimodal training and greatly enhances the system's perception robustness under fragmented data, but also takes into account the dual requirements of high precision and real-time performance of the low-altitude safety protection system with extremely low computational overhead.
[0014] This application solves the technical problems in the prior art, such as semantic conflicts caused by the awkward fusion of heterogeneous modal features and the inability to capture local features of small targets.
[0015] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the overall architecture of the algorithm network according to an embodiment of this application; Figure 2 A flowchart illustrating the method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the cross-scale adaptive feature enhancement architecture according to an embodiment of this application; Figure 4 This is a schematic diagram of the modal cue modulation architecture according to an embodiment of this application; Figure 5 This is a schematic diagram of the functional modules of an embodiment of the low-altitude micro-target detection device of this application; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] To address the shortcomings of existing technologies, refer to Figure 1 , Figure 1 This is a schematic diagram of the overall architecture of the algorithm network according to an embodiment of this application. Figure 1 As shown, the algorithm network architecture of this application mainly consists of a backbone network, a cross-scale adaptive feature enhancement module, a hybrid encoder, a cue modulation module, and a decoding prediction module. The modules are cascaded and connected in the order of feature flow to form a complete end-to-end detection link.
[0020] 1. Backbone network: A pre-trained PResNet-50 network is used to receive input visible light or infrared images and extract multi-scale hierarchical features; 2. Cross-scale adaptive feature enhancement module: pre-embedded after the multi-scale feature projection layer and before the hybrid encoder, used to pre-enhance features at three scales; 3. Hybrid encoder: Includes AIFITransformer (Attention-based Intra-scale Feature Interaction Transformer) and FPN, responsible for the interaction and aggregation of global features; 4. Cue Modulation Module: Located between the encoder and decoder, its function is to explicitly conditionally modulate the encoder output features according to the modality type of the current input image; 5. Transformer Decoder and Prediction Head: Receives refined features with modal context and performs the final prediction of the target class and bounding box.
[0021] It should be noted that the above network architecture and various algorithm modules can be mounted on computing devices with data processing capabilities (such as industrial control computers, edge computing boxes, or UAV-borne display and control terminals containing CPU, GPU / NPU processing chips and memory) in practical applications, forming a complete low-altitude dual-modal micro-target detection system in a hardware and software combination.
[0022] Furthermore, this application discloses a method for detecting small targets at low altitudes, comprising: Step S10: Through the backbone network, feature extraction is performed on the acquired image to be detected to obtain N hierarchical feature maps of different scales; In this embodiment, an image to be detected is acquired, wherein the image to be detected is an infrared or visible light image. The image to be detected is processed by a backbone network to extract three feature maps {P3, P4, P5} of different scales and levels, with spatial resolutions of [missing information - likely a percentage of the original image's resolution]. , and The backbone network is a pre-trained PRESNet-50 network (PaddleResidualNeural Network, a pre-trained residual neural network). The i-th feature map P3 represents the original image with the same spatial resolution. The high-resolution shallow feature map is rich in details such as edges and textures; the (i+1)th feature map P4 represents the original image with the same spatial resolution. The mid-resolution mid-layer feature map; the (i+2)th feature map P5 represents the spatial resolution of the original image. The low-resolution deep feature map is rich in high-level semantic information.
[0023] Step S20: Semantically refine the N feature maps to obtain a refined multi-scale feature map sequence; In this embodiment, the feature channels of the three scales are uniformly projected to 256 dimensions using a 1×1 convolution. The projected multi-scale features {P3, P4, P5} are then fed into the cross-scale adaptive feature enhancement module (CSAFE module) for semantic refinement.
[0024] Semantic refinement comprises two cascaded sub-stages: 1. Cross-scale adaptive fusion: A top-down, lightweight bottleneck interaction mechanism is used to pass high-resolution context. For the i-th feature map... First, adaptive average pooling is used to align the features down to the (i+1)th feature map. The spatial dimensions are then determined. Subsequently, semantic compression and expansion are performed using a bottleneck layer with a compression ratio of r=8 (containing C→C / 8 dimensionality reduction convolutions and C / 8→C dimensionality increase convolutions) to obtain the enhanced feature map.
[0025] The calculation formula is as follows:
[0026] Where σ represents the Sigmoid activation function; g represents the learnable gating scalar. In this embodiment, the learnable gating scalar is initialized to 0, and the degree of injection of shallow details into the deep semantic space is dynamically controlled through a soft gating mechanism. This represents the (i+1)th feature map after enhancement. This represents the (i+1)th feature map. Let i represent the i-th feature map, Bottleneck ( ) is the bottleneck network unit, AvgPool ( The downsampling operation is performed. This cascaded structure passes the data sequentially from P3 to P5, outputting an enhanced feature map sequence {P3, P4', P5'}. The i-th feature map P3 represents a high-resolution shallow feature map, the (i+1)-th enhanced feature map P4' represents a mid-level feature map enhanced by P3, and the (i+2)-th enhanced feature map P5' represents a high-level feature map enhanced by P4'.
[0027] A top-down, lightweight bottleneck interaction mechanism is introduced to overcome the limitation of traditional direct linear projection of multi-scale features, which easily leads to the loss of high-frequency details. By dynamically controlling the injection degree of shallow detail features into the deep semantic space through a learnable gating scalar (g), the local cues of small targets are effectively preserved and amplified during cross-layer downsampling.
[0028] 2. Scale-by-scale expert blending: The enhanced feature map sequence output by the cross-scale adaptive fusion module is {P3, P4', P5'}, where: P3: High-resolution shallow feature map ( × ); P4': Mid-layer feature map enhanced by P3 ( × ); P5': High-level feature map enhanced by P4' ( × ); H represents the height of the input image, and W represents the width of the input image.
[0029] The method of applying Scale-by-Scale Expert Hybrid (MoE) is to perform the same MoE processing flow independently on each of the three feature maps.
[0030] The specific steps are as follows: 1. Independent processing: P3, P4', and P5' are used as input feature maps, and each is fed into its corresponding MoE branch (e.g., ...). Figure 3 As shown, each of the three scales, P3', P4', and P5', is configured with an independent MoE structure, where the feature map... The feature map P is a tensor belonging to a B×C×H×W dimensional real space. Let B represent the set of real numbers. The superscripts B×C×H×W represent the dimensions and shape of the tensor. B represents the batch size, H represents the height of the input image, W represents the width of the input image, and C represents the number of channels of the feature map.
[0031] 2. Processing flow (taking P3 as an example): Perform global average pooling on P3 to obtain a global descriptor with dimensions [B, C, 1, 1].
[0032] By mapping through a two-layer fully connected network (C → C / 4 → 2) and normalizing using Softmax, the routing weights of the two experts are obtained. , Spatial and channel experts process P3 in parallel. Spatial experts perform 3×3 depthwise convolutions and 1×1 pointwise convolutions to refine local, minute target textures. Channel experts calculate channel weights using global pooling and two layers of 1×1 convolutions (dimensionality transformation from C→C / 16→C), and then recalibrate the channel dimensions using Sigmoid activation. Finally, the expert outputs are weighted and aggregated, and then superimposed onto the shallow feature map P3 via residual connections to obtain the refined shallow feature map. .
[0033] in, ; In the formula, This represents the refined shallow feature map; This represents the shallow feature map in the enhanced feature map sequence (i.e., the feature map output by the cross-scale adaptive fusion module that has not undergone expert processing). This represents the residual learning rate, initialized to 0.1, used to control the strength of the expert branch's correction of the original features, ensuring stable gradient backpropagation in the early stages of training. The routing weight of the spatial expert is the routing weight generated by the gating network for the spatial expert, which is normalized by Softmax and has a value range of [0,1]. The output feature map of the spatial expert branch is represented by extracting local spatial details through 3×3 depthwise convolution and 1×1 pointwise convolution. This represents the routing weights of the channel experts, where the routing weights of the channel experts are generated by the gating network, normalized by Softmax, and... , This represents the output feature map of the channel expert branch, which is recalibrated after calculating the channel weights through global pooling and two 1×1 convolutions.
[0034] 3. Parallel execution: P4' and P5' are processed in parallel in exactly the same way, and the three do not interfere with each other.
[0035] in, ; ; This represents the refined mid-layer feature map. This represents the refined high-level feature map.
[0036] In simple terms, the process is as follows: cross-scale fusion outputs three feature maps (P3: high-resolution shallow feature map; P4': mid-level feature map enhanced by P3; P5': high-level feature map enhanced by P4') → each feature map enters its own MoE branch → each feature map independently completes the process of "gating weight calculation → dual-expert parallel processing → weighted aggregation → residual output" → outputs three enhanced feature maps, i.e., the refined multi-scale feature map sequence. , , }
[0037] In this embodiment, a lightweight MoE architecture, including spatial experts (handling local texture boundaries) and channel experts (calibrating channel dimensions), is independently configured to address the different characteristics of multi-scale features. Routing weights are dynamically generated through a global gating network. , This enables adaptive weighted aggregation of spatial and channel domain features, greatly enhancing the discriminative feature representation of small targets.
[0038] Specifically, refer to Figure 3 , Figure 3 This is a schematic diagram of the cross-scale adaptive feature enhancement architecture according to an embodiment of this application, which is the core unit for solving the problem of feature loss in downsampling of small targets at low altitudes. Through a two-stage cascaded enhancement of cross-scale soft-gated fusion and scale-wise expert hybridization (MoE), the weak features of small targets are finely enhanced and semantic interference is suppressed.
[0039] In this embodiment, the input is the three-scale basic feature map output by the backbone network: P3[B,256,H / 8,W / 8]: High-resolution shallow feature map, rich in details such as small object edges and textures; P4[B,256,H / 16,W / 16]: Medium-resolution mid-level feature map; P5[B,256,H / 32,W / 32]: Low-resolution deep feature map, rich in high-level semantic information.
[0040] The output is a refined sequence of multi-scale feature maps. , , The data is then fed into the subsequent AIFI Transformer and FPN structure, and finally input into the decoder to complete target detection.
[0041] Level 1: Soft-gated cross-scale fusion submodule: like Figure 3 As shown in the dashed box, the soft-gated cross-scale fusion submodule adopts a top-down interaction mechanism to achieve controllable injection of shallow details into the deep semantic space, thus solving the problem of loss of details of small targets.
[0042] Taking the interaction from P3 to P4 as an example, scale alignment and feature compression are performed on the shallow feature map P3 by adaptive average pooling to align its spatial size to the resolution (H / 16, W / 16) of the deep feature map P4. Then, semantic compression and reconstruction are completed through 1×1 convolution (i.e., bottleneck layer structure, compression ratio r=8), which reduces the amount of computation while preserving key details.
[0043] Learnable soft-gated injected aligned and compressed features are dynamically weighted by a learnable gated scalar g (initialized to 0) through a sigmoid activation function σ(g), and then weighted and fused with the deep feature map P4. This design enables adaptive injection of shallow details into deep features, avoiding detail redundancy or loss caused by fixed-weight fusion. Similarly, P4' and P5 repeat the above interaction to obtain the enhanced feature map P5', completing the cross-scale fusion stage.
[0044] Level 2: Scale-by-Scale Expert Hybrid (MoE) Submodule: like Figure 3As shown below, independent MoE branches are configured for the enhanced feature maps P3, P4', and P5' output by the cross-scale adaptive fusion module, respectively, to achieve refined enhancement of the features of small targets through dual heterogeneous experts.
[0045] Taking the P3 scale feature map as an example, the processing flow is as follows: The gating network generates routing weights by first compressing the input features through global average pooling, then mapping them through two fully connected layers (dimensionality transformation C→C / 4→2), and finally normalizing them using Softmax to obtain the routing weights from two experts. (Space expert) and (Channel expert).
[0046] Parallel processing of two heterogeneous experts: Spatial Expert: Employing a lightweight structure of 3×3 depthwise convolution and 1×1 pointwise convolution, it focuses on extracting high-frequency spatial details such as edges and textures of small targets, and enhances local weak features.
[0047] Channel expert: Employs a structure of global pooling and two 1×1 convolutions (dimensional transformation C→C / 16→C) to calculate the importance weights of channel dimensions, suppress redundant channel interference caused by modal heterogeneity, and calibrate feature distribution.
[0048] Residual aggregation output enhancement features: The output feature maps of the two experts are weighted and aggregated according to the routing weights, and then superimposed onto the original input features through residual connections. .
[0049] The feature map processing for P4' and P5' scales is the same as described above, and the final output is... and The two-level enhancement of the cross-scale adaptive feature enhancement module (CSAFE module) is completed.
[0050] The refined multi-scale feature map sequence output by the Cross-Scale Adaptive Feature Enhancement Module (CSAFE module) { , , The data is fed into the AIFI Transformer to capture global context associations, then aggregated from top to bottom using the FPN structure, and finally sent to the Transformer decoder.
[0051] This embodiment effectively solves the problem of losing details of small targets during standard downsampling by combining soft-gated cross-scale fusion with scale-by-scale expert hybridization. At the same time, it suppresses feature interference caused by modal heterogeneity, providing high-quality features with both detail and semantic consistency for subsequent detection tasks.
[0052] Step S30: Based on the refined multi-scale feature map sequence, multi-scale feature aggregation is performed through a hybrid encoder to obtain an aggregated multi-scale feature map sequence. In this embodiment, the multi-scale feature map sequence refined by CSAFE { , , The features are fed into a hybrid encoder. The highest-level semantic features are fed into the AIFI Transformer to capture global associations, and then the overall features are fed into the FPN for top-down multi-scale feature aggregation, resulting in a sequence of aggregated multi-scale feature maps. The AIFI Transformer is a lightweight Transformer encoder that performs self-attention only on the highest-level feature maps.
[0053] Specifically, in one embodiment, global association modeling (AIFI Transformer) is performed on the highest-level semantic features: First, select the highest-level feature map in the sequence that has the richest semantic information. The data is then fed into the AIFITransformer module for global correlation modeling.
[0054] In this module: Two-dimensional spatial feature map The flattened form is a one-dimensional sequence, and positional encoding is introduced to preserve spatial location information; By using the self-attention mechanism of Transformer, long-distance dependencies between global features are modeled to capture the contextual relationships between small targets and background and interference objects in low-altitude scenes. This step injects global semantic consistency constraints into the highest-level features without significantly increasing computational overhead, thereby improving the model's robustness in recognizing weak targets and targets with low signal-to-noise ratios.
[0055] After modeling is completed, an enhanced global semantic feature map is output. With the original , Together, they constitute the complete feature set of the hybrid encoder.
[0056] For top-down multi-scale feature fusion aggregation (FPN): The feature set enhanced by AIFI Transformer is fed into the Feature Pyramid Network (FPN) to perform top-down multi-scale feature fusion aggregation.
[0057] Top-down upsampling and fusion using the highest-level global semantic feature map Starting from this point, the lower-level feature maps are sequentially upsampled and fused: right Perform upsampling to make its spatial resolution equal to Alignment; The upsampled high-level semantic information is injected through element-wise addition or concatenation. , to obtain the fused feature map ; Repeat the above process to Upsampling and The feature maps are then fused to obtain the final multi-scale fused feature map. .
[0058] After being aggregated by FPN, the multi-scale feature map outputs produce a sequence of aggregated multi-scale feature maps that combine high-level global semantics with low-level detailed information. , , This feature set also includes: High-resolution feature maps rich in minute target details; : A transitional feature map with medium resolution, detail, and semantic balance; Low-resolution, globally semantically enhanced high-level feature maps.
[0059] The hybrid encoder in this embodiment achieves global context modeling of the highest-level features through the combination of "AIFI Transformer and FPN", which makes up for the lack of global correlation information in the traditional CNN downsampling process; Top-down multi-scale feature aggregation transmits high-level global semantics to low-level features, strengthening weak semantic cues of small targets. The output multi-scale fusion features, when fed into the subsequent FiLM modal cueing modulation module, can simultaneously meet the dual requirements of "detail richness" and "semantic consistency", effectively improving the detection accuracy of small targets under dual-modal conditions.
[0060] Step S40: Based on the aggregated multi-scale feature map sequence and the discrete modal identifier corresponding to the image to be detected, modal conditional modulation processing is performed to obtain a modal-aware refined feature sequence. In this embodiment, before heterogeneous data enters the decoder, the network first performs modal conditionalization processing at the encoder output: 1. Modal embedding: Define a learnable embedding matrix Wherein, the embedding matrix E is a 2×64 real matrix. Representation matrix. Identify the discrete modes corresponding to the input image to be detected. Mapped to continuous mode vectors Where 0 represents visible light and 1 represents infrared light.
[0061] 2. Shared parameter generation: based on continuous mode vectors A single shared MLP (Multi-Layer Perceptron) network is used to sequentially perform Linear (64→128), SiLU activation, and Linear (128→1536) operations, generating a parameter tensor with a total dimension of B×1536. To avoid conditional injection disrupting the high-quality pre-trained feature distribution, the weights and biases of the last layer of this MLP are forced to be zero-initialized. Forcing the zero-initialization of the weights and biases of the last layer of the shared MLP ensures that, in the early stages of network training, the modulation operation strictly degenerates into an identity mapping, effectively preventing conditional injection from damaging the high-quality pre-trained feature distribution and guaranteeing a smooth transition and high robustness of model training.
[0062] 3. Feature-level linear modulation: The total parameter tensor is divided according to the number of channels in each layer, and the channel-level scaling coefficients for each layer's features are extracted. With translation coefficient ,in, Represents a vector. This represents the vector dimension. Subsequently, feature-level linear modulation is performed along the channel dimension to obtain the modality-aware refined feature sequence { , , }
[0063] ; ; ; Here, ⊙ represents element-wise multiplication. This represents shallow refined features of modality perception. This indicates the refined features of the shallow to mid-level modal perception layer. This represents high-level refined features of modality perception; express Channel-level scaling factor, express Translation coefficient; express Channel-level scaling factor, express Translation coefficient; express Channel-level scaling factor, express The translation coefficients. Due to the presence of zero initialization, the modulation in the early stages of training strictly degenerates into an identity mapping ( , , As training progresses, the network smoothly learns modality-specific conditional biases.
[0064] Specifically, refer to Figure 4 , Figure 4 This is a schematic diagram of the modal cue modulation architecture according to an embodiment of this application, which is used to achieve adaptive conditional correction of dual-modal input features, eliminate semantic conflicts between infrared and visible light heterogeneous data, and provide refined features for modal adaptation for the subsequent Transformer decoder.
[0065] like Figure 4 As shown, the inputs to this module include: The multi-scale feature map sequence of the aggregated visible light image (modal identifier id=0, corresponding to the visible light (RGB) image) and the multi-scale feature map sequence of the aggregated infrared image (modal identifier id=1, corresponding to the infrared light (IR) image). First, modal embedding is performed, mapping the input discrete modal identifiers (0 or 1) to continuous modal vectors p with dimensions [B, 64] through a learnable embedding matrix. Here, B represents the batch size and 64 represents the embedding dimension. Its purpose is to transform binary modal information into continuous features that the network can understand, serving as subsequent conditional control signals.
[0066] The generated mode vector p is fed into a single shared MLP network. The network structure is as follows: First layer: Linear transformation (64 → 128); Activation layer: SiLU activation function; The second layer is a linear transformation, Linear(128 → 1536), where the parameters are forced to be zero-initialized. This MLP ultimately outputs a total parameter tensor with dimensions [B, 1536]. The zero-initialization design ensures that the modulation operation in the early stages of training is equivalent to an identity mapping, avoiding disruption of the pre-trained feature distribution.
[0067] Feature-level linear modulation (FiLM): The generated [B, 1536] parameter tensor is divided according to the number of channels of each layer of features, resulting in an aggregated multi-scale feature map sequence. , , Generate the corresponding scaling factors respectively. Translation coefficient The P3 dimension is [B, 256, H / 8, W / 8]; the P4 dimension is [B, 256, H / 16, W / 16]; and the P5 dimension is [B, 256, H / 32, W / 32]. Subsequently, FiLM modulation is performed on the feature map at each scale: ; ; ; Through this operation, the model can dynamically adjust the feature distribution according to the input mode, so that the infrared and visible light features have completed adaptive calibration before entering the decoder.
[0068] After modulation, the aggregated multi-scale feature map sequence is converted into a modality-aware refined feature sequence. The refining characteristics at each level shown in the diagram. These features are ultimately fed into the Transformer decoder on the right for subsequent object detection prediction.
[0069] In this embodiment, by introducing modal embedding, discrete modal identifiers are mapped to continuous conditional vectors. A shared network is used to generate channel-level scaling coefficients (γ) and translation coefficients (β). Linear affine transformations are performed on multi-scale features, giving the decoder explicit modal adaptive perception and decoupling capabilities. This achieves a lightweight, stable, and efficient modal conditional modulation mechanism, solving the problem that direct fusion of infrared and visible light heterogeneous data in traditional dual-modal detection methods can easily lead to semantic conflicts.
[0070] Step S50: Input the modality-aware refined features into the decoder to obtain the class score and bounding box position coordinates of the low-altitude small target output by the decoder.
[0071] In this embodiment, refined features with explicit modal context are used. The data is fed into the Transformer decoder. The decoder uses a cross-attention mechanism to search and match the target query in the refined bimodal feature map, and finally outputs the accurate category score and bounding box position of the low-altitude small target (drone / bird).
[0072] For example, suppose the input is a visible light image containing a distant drone (approximately 15×15 pixels). A target query in the decoder is retrieved through cross-attention. The region where the drone is located (highest response value) is activated in the (high-resolution feature map). Through self-attention interaction, the repeated prediction of the target by other queries is suppressed. Finally, the classification head outputs the category score of "drone" with 0.92, and the regression head outputs the bounding box coordinates [x=320, y=240, W=32, H=32].
[0073] In this embodiment, on the one hand, by using a cross-scale adaptive feature enhancement module and a scale-wise hybrid expert mechanism, high-resolution local details are actively injected into the deep semantic space for semantic refinement. This effectively overcomes the problem that traditional networks are prone to losing weak high-frequency features during downsampling, and significantly improves the feature retention rate and final detection accuracy of small targets (such as long-distance drones and birds) in complex low-altitude environments.
[0074] On the other hand, by adopting a modal conditional modulation mechanism based on zero initialization, adaptive perception and feature decoupling of the input modality are achieved, effectively eliminating semantic conflicts and domain shifts caused by the fusion of heterogeneous infrared and visible light data. Combined with the overall minimalist and lightweight design, it not only solves the feature oscillation problem in multimodal training and greatly enhances the system's perception robustness under fragmented data, but also takes into account the dual requirements of high precision and real-time performance of the low-altitude safety protection system with extremely low computational overhead.
[0075] This embodiment solves the technical problems in the prior art, such as semantic conflicts caused by the awkward fusion of heterogeneous modal features, the inability to capture local features of small targets, and the lack of adaptive modal perception and dynamic processing capabilities of the model.
[0076] Based on the same inventive concept, embodiments of this application also provide a low-altitude small target detection device.
[0077] In one embodiment, reference is made to Figure 5 , Figure 5 This is a functional module diagram of an embodiment of the low-altitude small target detection device of this application. Figure 5 As shown, the low-altitude small target detection device includes: The feature extraction module 10 is configured to extract features from the acquired image to be detected through the backbone network to obtain N feature maps of different scales, where N is a positive integer; Feature enhancement module 20 is configured to perform semantic refinement on N feature maps to obtain a refined multi-scale feature map sequence; The hybrid encoding module 30 is configured to perform multi-scale feature aggregation based on the refined multi-scale feature map sequence using a hybrid encoder to obtain an aggregated multi-scale feature map sequence. The modulation module 40 is configured to perform modality conditional modulation processing based on the aggregated multi-scale feature sequence and the discrete modality identifier corresponding to the image to be detected, to obtain modality-aware refined features. The feature decoding module 50 is configured to input the modality-aware refined features into the decoder to obtain the class score and bounding box position coordinates of the low-altitude small target output by the decoder.
[0078] Optionally, in one embodiment, the feature enhancement module 20 is configured to: Each feature map is semantically compressed and expanded to obtain an enhanced multi-scale feature map sequence. The outputs of spatial experts and channel experts are aggregated by weight and then superimposed onto the enhanced feature map sequence through residual connections to obtain a refined multi-scale feature map sequence.
[0079] Optionally, in one embodiment, the feature enhancement module 20 is further configured to: Each feature map is input into the cross-scale adaptive fusion module, and semantic compression and expansion are performed on each feature map using a first preset formula to obtain each enhanced feature map; The first preset formula is as follows:
[0080] In the formula, σ represents the Sigmoid activation function; g represents the learnable gating scalar. This represents the (i+1)th feature map after enhancement. This represents the feature map of the (i+1)th image. Let Bouttleneck represent the feature map of the i-th image. ) is the bottleneck network unit, AvgPool ( () represents a downsampling operation; Based on each enhanced feature map, an enhanced multi-scale feature map sequence is constructed.
[0081] Optionally, in one embodiment, the feature enhancement module 20 is further configured to: Obtain the routing weights of spatial experts, the feature maps output by spatial experts, the routing weights of channel experts, and the feature maps output by channel experts; Based on each enhanced feature map in the enhanced feature map sequence, the routing weight of the spatial expert, the feature map output by the spatial expert, the routing weight of the channel expert, and the feature map output by the channel expert, each refined feature map is calculated using the second preset formula. The second preset formula is as follows:
[0082] In the formula, This represents the refined feature map; This represents the enhanced feature map. Represents the residual learning rate. This represents the routing weight of the spatial expert. This represents the feature map output by the space expert. This represents the routing weight of the channel expert. This represents the feature map output by the channel expert; Based on each refined feature map, a refined multi-scale feature map sequence is constructed.
[0083] Optionally, in one embodiment, the hybrid encoding module 30 is configured to: Based on the highest-level feature map in the refined multi-scale feature map sequence, global association modeling is performed to obtain an enhanced global semantic feature map. Based on the enhanced global semantic feature map and the other layer feature maps in the refined multi-scale feature map sequence, a top-down multi-scale feature fusion and aggregation is performed through a feature pyramid network to obtain the aggregated multi-scale feature map sequence.
[0084] Optionally, in one embodiment, the prompt modulation module 40 is configured to: The discrete modal identifiers corresponding to the image to be detected are mapped to continuous modal vectors; The modality vector is input into a shared multilayer perceptron network to obtain the total parameter tensor output by the shared multilayer perceptron network. The total parameter tensor is divided according to the number of channels of each layer of features to obtain the channel-level scaling coefficient and translation coefficient corresponding to each feature map in the aggregated multi-scale feature map sequence; Based on each feature map in the aggregated multi-scale feature map sequence, the channel-level scaling factor, and the translation factor, feature-level linear modulation is performed in the channel dimension to obtain a modality-aware refined feature sequence.
[0085] Optionally, in one embodiment, the shared multilayer perceptron network includes a first linear transformation layer, an activation layer, and a second linear transformation layer.
[0086] The functions of each module in the aforementioned low-altitude small target detection device correspond to the steps in the aforementioned low-altitude small target detection method embodiment, and their functions and implementation processes will not be described in detail here.
[0087] Based on the same inventive concept, embodiments of this application also provide an electronic device, the structure of which is as follows: Figure 6As shown, it includes: a memory and a processor, wherein the processor is used to read and execute the computer program stored in the memory to implement the aforementioned method for detecting small targets at low altitudes.
[0088] Based on the same inventive concept, this application also provides a computer storage medium storing computer-executable instructions, which, when executed, implement the aforementioned low-altitude small target detection method.
[0089] Finally, it should be noted that while some processes described in the embodiments of this application include multiple operations or steps that appear in a specific order, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.
[0090] Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for detecting small targets at low altitudes, characterized in that, The method includes: The backbone network is used to extract features from the acquired image to be detected, resulting in N feature maps of different scales, where N is a positive integer. Semantic refinement is performed on the N feature maps to obtain a refined multi-scale feature map sequence. Based on the refined multi-scale feature map sequence, multi-scale feature aggregation is performed through a hybrid encoder to obtain an aggregated multi-scale feature map sequence. Based on the aggregated multi-scale feature sequence and the discrete modal identifier corresponding to the image to be detected, modal conditional modulation processing is performed to obtain modal-aware refined features. The modality-aware refined features are input into the decoder to obtain the class score and bounding box position coordinates of the low-altitude small target output by the decoder.
2. The low-altitude small target detection method according to claim 1, characterized in that, The semantic refinement of the N feature maps to obtain a refined multi-scale feature map sequence includes: Each feature map is semantically compressed and expanded to obtain an enhanced multi-scale feature map sequence. The outputs of spatial experts and channel experts are aggregated by weight and then superimposed onto the enhanced feature map sequence through residual connections to obtain a refined multi-scale feature map sequence.
3. The low-altitude small target detection method according to claim 2, characterized in that, The step of semantically compressing and expanding each feature map to obtain an enhanced multi-scale feature map sequence includes: Each feature map is input into the cross-scale adaptive fusion module, and semantic compression and expansion are performed on each feature map using a first preset formula to obtain each enhanced feature map; The first preset formula is as follows: In the formula, σ represents the Sigmoid activation function; g represents the learnable gating scalar. This represents the (i+1)th feature map after enhancement. This represents the feature map of the (i+1)th image. Let Bouttleneck represent the feature map of the i-th image. ) is the bottleneck network unit, AvgPool ( () represents a downsampling operation; Based on each enhanced feature map, an enhanced multi-scale feature map sequence is constructed.
4. The low-altitude small target detection method according to claim 2, characterized in that, The process of weighted aggregation of the outputs of spatial experts and channel experts, and superimposing them onto the enhanced feature map sequence via residual connections, yields a refined multi-scale feature map sequence, including: Obtain the routing weights of spatial experts, the feature maps output by spatial experts, the routing weights of channel experts, and the feature maps output by channel experts; Based on each enhanced feature map in the enhanced feature map sequence, the routing weight of the spatial expert, the feature map output by the spatial expert, the routing weight of the channel expert, and the feature map output by the channel expert, each refined feature map is calculated using the second preset formula. The second preset formula is as follows: In the formula, This represents the refined feature map; This represents the enhanced feature map. Represents the residual learning rate. This represents the routing weight of the spatial expert. This represents the feature map output by the space expert. This represents the routing weight of the channel expert. This represents the feature map output by the channel expert; Based on each refined feature map, a refined multi-scale feature map sequence is constructed.
5. The method for detecting low-altitude small targets according to claim 1, characterized in that, The refined multi-scale feature map sequence is used to perform multi-scale feature aggregation through a hybrid encoder to obtain an aggregated multi-scale feature map sequence, including: Based on the highest-level feature map in the refined multi-scale feature map sequence, global association modeling is performed to obtain an enhanced global semantic feature map. Based on the enhanced global semantic feature map and the other layer feature maps in the refined multi-scale feature map sequence, a top-down multi-scale feature fusion and aggregation is performed through a feature pyramid network to obtain the aggregated multi-scale feature map sequence.
6. The method for detecting low-altitude small targets according to claim 1, characterized in that, The modality conditional modulation process is performed based on the aggregated multi-scale feature map sequence and the discrete modality identifier corresponding to the image to be detected to obtain a modality-aware refined feature sequence, including: The discrete modal identifiers corresponding to the image to be detected are mapped to continuous modal vectors; The modality vector is input into a shared multilayer perceptron network to obtain the total parameter tensor output by the shared multilayer perceptron network. The total parameter tensor is divided according to the number of channels of each layer of features to obtain the channel-level scaling coefficient and translation coefficient corresponding to each feature map in the aggregated multi-scale feature map sequence; Based on each feature map in the aggregated multi-scale feature map sequence, the channel-level scaling factor, and the translation factor, feature-level linear modulation is performed in the channel dimension to obtain a modality-aware refined feature sequence.
7. The method for detecting low-altitude small targets according to claim 6, characterized in that, The shared multilayer perceptron network includes a first linear transformation layer, an activation layer, and a second linear transformation layer.
8. A low-altitude micro-target detection device, characterized in that, The device includes: The feature extraction module is configured to extract features from the acquired image to be detected through the backbone network, and obtain N feature maps of different scales, where N is a positive integer; The feature enhancement module is configured to perform semantic refinement on the N feature maps to obtain a refined multi-scale feature map sequence. The hybrid encoding module is configured to perform multi-scale feature aggregation based on the refined multi-scale feature map sequence using a hybrid encoder to obtain an aggregated multi-scale feature map sequence. The prompt modulation module is configured to perform modality conditional modulation processing based on the aggregated multi-scale feature sequence and the discrete modality identifier corresponding to the image to be detected, to obtain modality-aware refined features; The feature decoding module is configured to input the modality-aware refined features into the decoder to obtain the class score and bounding box position coordinates of the low-altitude small target output by the decoder.
9. An electronic device, characterized in that, include: Memory, processor; The processor is configured to read and execute the computer program stored in the memory to implement the steps of the low-altitude small target detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, implement the steps of the low-altitude small target detection method according to any one of claims 1-7.