A lightweight small target detection method based on calibration driving

CN122821084APending Publication Date: 2026-09-25UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610871712.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

一是小目标细节增强缺乏选择性,容易同时放大目标结构和背景高频干扰,导致误检增加;二是现有跨尺度特征融合多为单次语义传播,难以持续校准高分辨率特征中的小目标语义响应;三是现有边界框回归损失在小目标低重叠、无重叠或包含关系下容易出现中心校正梯度不足,影响定位精度;四是现有轻量化检测器受限于模型容量,若单纯增加复杂模块又会破坏其部署效率

Benefits of technology

1、提高小目标频率细节表达的有效性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821084A_ABST
    Figure CN122821084A_ABST
Patent Text Reader

Abstract

The application discloses a kind of light-weight small target detection methods based on calibration drive, belong to computer vision and target detection technical field, this light-weight small target detection method includes: obtaining image to be detected;The image to be detected obtained is input into the preset light-weight small target detection model;Wherein, the light-weight small target detection model includes backbone network, neck network and detection head;And the light-weight small target detection model introduces calibration mechanism, to realize the calibration at least one direction in feature representation, feature fusion and boundary box regression gradient in small target detection;Detection result is output using light-weight small target detection model.The present application can not rely on large-scale model expansion, under the premise of maintaining lower parameter amount and calculation amount, while frequency detail representation, cross-scale semantic fusion and boundary box regression supervision are improved in a targeted manner, so as to improve the small target detection precision, positioning stability and deployment applicability under complex background and resource limited conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and target detection technology, and in particular to a lightweight small target detection method based on calibration-driven methods. Background Technology

[0002] Small object detection aims to locate and identify small objects with a low pixel count in images or videos. Compared to regular object detection, small objects typically occupy only a small number of pixels in an image, and their edge, contour, texture, and positional information is weak, making them easily weakened or lost during feature extraction and downsampling. Therefore, small object detection places higher demands on the model's detail perception, cross-scale feature fusion, and bounding box localization capabilities. For example, datasets such as VisDrone and AI-TOD are designed for small object detection scenarios in drone aerial photography or remote sensing images, where the objects are generally characterized by small scale, dense distribution, and complex backgrounds.

[0003] Small target detection has significant application value in scenarios such as drone aerial photography, remote sensing image analysis, intelligent transportation, security monitoring, robot perception, and edge vision perception. For example, in drone aerial images, targets such as vehicles, pedestrians, and cyclists often exhibit characteristics such as small scale, dense distribution, severe occlusion, and complex backgrounds due to factors such as long shooting distance, large field of view, and limited imaging resolution. At the same time, background content such as road markings, building textures, shadows, vegetation, and sensor noise can also produce local texture responses similar to small targets, leading to problems such as missed detections, false detections, and inaccurate localization by the detector.

[0004] Deep learning-based object detection methods typically consist of a backbone network, a neck network, and a detection head. The backbone network extracts image features, the neck network fuses features at different scales, and the detection head performs object classification, object confidence prediction, and bounding box regression. For small object detection, high-resolution shallow features preserve more spatial details, which is beneficial for object localization; low-resolution deep features have stronger semantic expressive power, which is beneficial for category discrimination. Therefore, how to effectively fuse shallow detail information and deep semantic information is an important technical problem in small object detection. Specifically, for small object detection tasks, due to the small object size and low pixel ratio, detection performance mainly depends on the stability of detail feature representation, cross-scale semantic fusion, and bounding box regression supervision.

[0005] Existing feature fusion methods typically employ feature pyramid structures. For example, FPN (Feature Pyramid Network) transmits semantic information to high-resolution features through a top-down path, PAN (Path Aggregation Network) enhances multi-scale feature aggregation through a bottom-up path, and BiFPN (Bi-directional Feature Pyramid Network) further utilizes learnable weights for multi-layer feature fusion.

[0006] On the other hand, in practical applications, small target detection models typically need to be deployed on drones, vehicle-mounted devices, embedded platforms, mobile terminals, or edge computing devices. These devices are often limited by factors such as computing power, storage space, power consumption, and real-time requirements. Therefore, the detection model not only needs to have high detection accuracy but also needs to maintain a low number of parameters and computational cost. Lightweight detectors have thus become an important technical approach for the practical deployment of small target detection. Existing lightweight or small target detection methods typically improve detection performance by enhancing high-resolution features, introducing multi-scale feature fusion structures, designing lightweight modules, or improving bounding box regression losses. For example, methods such as CFPT (Cross-Scale Feature Pyramid Transformer), FBRT-YOLO-X, and CAEM-DETR all attempt to improve the performance of small target detection.

[0007] In feature representation enhancement, existing methods typically improve the response of small targets by preserving high-resolution features, enhancing shallow details, introducing attention mechanisms, or modeling frequency domain information. The basic principle of these methods is to enhance details such as target edges, corners, contours, and local textures, making the small target appear more prominent in the feature map, thus facilitating classification and localization by subsequent detection heads. However, high-frequency details in aerial and remote sensing images do not originate solely from the small target itself. Background areas such as road markings, building roofs, shadows, vegetation textures, and sensor noise also generate significant high-frequency responses. If existing methods only directly enhance shallow details or high-frequency information without distinguishing between target-related high-frequency information and background interference high-frequency information, they can easily amplify both the small target structure and background clutter simultaneously, causing background areas to produce target-like responses, thereby increasing the risk of false detection. Therefore, existing feature enhancement methods still lack a mechanism for selectively calibrating frequency details.

[0008] In cross-scale feature fusion, existing object detectors typically employ a feature pyramid structure, transferring semantic information from low-resolution deep features to high-resolution shallow features. This type of structure usually includes a top-down path, lateral connections, upsampling operations, and multi-scale detection heads. Its working principle is to utilize deep features to provide category semantic information and shallow features to provide spatial detail information, fusing the two for the detection of objects at different scales. However, existing feature pyramid structures mostly belong to a single semantic propagation mechanism. Deep semantic information is used for detection and prediction after a single upsampling and lateral connection, lacking continuous feedback and recalibration of the fusion result. For small targets, the target region occupies only a small number of pixels, and deep semantics are easily diluted when transferred to high-resolution features, making it difficult to concentrate on the true small target region. Simultaneously, high-resolution shallow features contain a large amount of background texture, which can easily interfere with the semantic response of small targets. Therefore, existing cross-scale feature fusion methods are prone to problems such as scattered semantic responses, cross-scale semantic drift, and false background detections.

[0009] In bounding box regression optimization, existing object detectors typically employ loss functions based on Cross-Union Ratio (CUI) or its variants to constrain predicted bounding boxes. These loss functions guide the predicted bounding box to gradually approach the ground truth bounding box by measuring the degree of overlap, center distance, circumscribed region, or shape difference between the predicted and ground truth bounding boxes. For objects of normal size, a relatively stable overlap relationship usually forms between the predicted and ground truth bounding boxes, thus this type of regression supervision can effectively guide bounding box localization. However, for small objects, the bounding box size is small, and even a small pixel shift in the center of the predicted bounding box can lead to a significant decrease in CUI, or even cause the predicted and ground truth bounding boxes to completely lose overlap. In the early stages of training or when the localization deviation is large, regression loss based on overlapping regions is prone to insufficient gradient, making it difficult to effectively push the center of the predicted bounding box to quickly approach the ground truth object. Furthermore, when the predicted and ground truth bounding boxes are in a containment relationship or have low overlap, changes in CUI are not sensitive enough to center shifts, which also weakens the center correction capability. Therefore, existing bounding box regression methods still suffer from unstable regression gradients, slow convergence speed, and insufficient localization accuracy in small object detection.

[0010] In lightweight detectors, existing methods typically reduce parameter and computational costs by decreasing network depth, narrowing channel width, employing lightweight convolutions, simplifying the neck network, or compressing the detection head. These methods are suitable for resource-constrained scenarios such as drones, embedded devices, mobile terminals, and edge computing platforms, and can meet the requirements of real-time inference and low-power deployment to some extent. However, due to limited network capacity, lightweight models generally have weaker detailed feature representation capabilities, cross-scale semantic fusion capabilities, and bounding box regression capabilities compared to large detectors. Further introducing complex attention modules, multi-branch structures, or large fusion networks to improve small object detection accuracy significantly increases model parameter and computational costs, weakening the advantages of lightweight deployment. Therefore, existing lightweight small object detectors still struggle to achieve a good balance between detection accuracy and deployment efficiency.

[0011] In summary, the existing technology has the following problems: First, the enhancement of small target details lacks selectivity and easily amplifies both the target structure and high-frequency background interference, leading to an increase in false detections. Second, existing cross-scale feature fusion is mostly a single semantic propagation, making it difficult to continuously calibrate the semantic response of small targets in high-resolution features. Third, existing bounding box regression loss is prone to insufficient center correction gradients when small targets have low overlap, no overlap, or containment relationships, affecting localization accuracy. Fourth, existing lightweight detectors are limited by model capacity, and simply adding complex modules would compromise their deployment efficiency. Summary of the Invention

[0012] This invention provides a lightweight small target detection method based on calibration-driven approach, which at least partially solves the aforementioned technical problems existing in the prior art.

[0013] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On one hand, the present invention provides a lightweight small target detection method based on calibration-driven methods, including: Acquire the image to be detected; The acquired image to be detected is input into a preset lightweight small object detection model; wherein, the lightweight small object detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract image features, the neck network is used to fuse features at different scales; the detection head is used to complete object classification, object confidence prediction, and bounding box regression; and the lightweight small object detection model introduces a calibration mechanism to achieve calibration of at least one direction of the feature representation, feature fusion, and bounding box regression gradient in small object detection; The detection results are output using the lightweight small target detection model.

[0014] Furthermore, the backbone network includes multiple feature extraction stages, and a frequency calibration module is set in one or more feature extraction stages. After the image to be detected is input into the backbone network, it undergoes convolution, normalization, activation, and downsampling processing to obtain basic feature maps of different scales. In the feature extraction stage with the frequency calibration module, the basic feature maps are input to the frequency calibration module for frequency calibration to obtain frequency-calibrated output features. The frequency calibration process of the frequency calibration module includes: For the feature map of the input frequency calibration module, discrete wavelet transform is used to decompose it into subbands to obtain low-frequency subbands and multiple high-frequency subbands. For the low-frequency subband, a lightweight channel gating structure is used to generate low-frequency calibration weights, and the low-frequency subband is adjusted using the low-frequency calibration weights to obtain the calibrated low-frequency subband. For the high-frequency subband, a lightweight convolutional thinning structure is first used to refine the local features, and then high-frequency calibration weights are used to selectively enhance or suppress the high-frequency subband to obtain the calibrated high-frequency subband. The calibrated low-frequency subband and high-frequency subband are subjected to inverse wavelet transform to reconstruct the spatial domain frequency calibration residual signal. The spatial domain frequency calibration residual signal is residually fused with the feature map of the input frequency calibration module to obtain the frequency-calibrated output feature.

[0015] Furthermore, the lightweight channel gating structure includes a global average pooling layer, a multilayer perceptron, a nonlinear activation function, and a sigmoid function.

[0016] Furthermore, the lightweight convolutional thinning structure includes depthwise convolution and pointwise convolution.

[0017] Furthermore, the neck network includes a closed-loop cross-scale semantic calibration mechanism; The process by which the neck network fuses features at different scales includes: The system receives multi-scale features output from the backbone network, performs channel alignment processing on features of different scales, and gives each scale feature a preset channel dimension to form an initial pyramid feature state. The initial pyramid feature state includes a high-resolution feature state, an intermediate-resolution feature state, and a low-resolution semantic feature state. Multiple rounds of closed-loop semantic calibration are performed on the initial pyramid feature state. In each round of calibration, the high-resolution feature state, intermediate-resolution feature state, and low-resolution semantic feature state from the previous round are first scale-aligned. The scale-aligned multi-source features are then adaptively fused to obtain the high-resolution calibration features for the current round. Subsequently, based on the high-resolution calibration features of the current round, the pyramid feature states at each scale are updated through upsampling, downsampling, and lateral connections. The updated feature states are then input into the next round of calibration. After multiple rounds of closed-loop semantic calibration, the fusion of features at different scales is achieved.

[0018] Furthermore, the adaptive fusion of scale-aligned multi-source features includes: Channel-level fusion weights are generated based on global contextual information of features at different scales; among them, the more complex the image background texture, the greater the fusion weight of the low-resolution semantic feature state; the weaker the details of the small target location, the greater the fusion weight of the high-resolution feature state. The channel-level fusion weights are used to perform weighted aggregation of the scale-aligned high-resolution feature states, intermediate-resolution feature states, and low-resolution semantic feature states.

[0019] Furthermore, the scale alignment is achieved through upsampling, downsampling, or interpolation operations.

[0020] Furthermore, the bounding box regression optimization objective of the lightweight small target detection model during the training phase is a weighted combination of bounding box regression loss and gradient calibration loss; wherein, the gradient calibration loss includes a regression state scheduling factor and a center calibration potential function; The regression state scheduling factor dynamically adjusts the weight of the center correction term based on the overlap quality between the predicted and ground truth boxes; the lower the overlap quality between the predicted and ground truth boxes, the greater the weight of the center correction term. The center calibration potential function generates a center offset penalty based on the distance between the center of the predicted box and the center of the ground truth box, and normalizes it according to the scale of the ground truth box, so that small targets can obtain a stronger center correction gradient under the same pixel offset.

[0021] In another aspect, the present invention also provides an electronic device comprising a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described method.

[0022] In another aspect, the present invention also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the above method.

[0023] The beneficial effects of the technical solution provided by this invention include at least the following: 1. Improve the effectiveness of expressing the frequency details of small targets. This invention decomposes the input features into low-frequency and high-frequency sub-bands using a frequency calibration module and performs adaptive calibration on different frequency components. This enables the detection network to enhance edge, contour, and fine-grained texture information related to small targets, while suppressing high-frequency background interference such as road markings, building textures, shadows, vegetation textures, and sensor noise. This alleviates the problem of increased false detections caused by the blind enhancement of high-frequency information in existing methods.

[0024] 2. Improve the semantic dilution problem in high-resolution features

[0025] This invention uses a semantically calibrated neck network to perform closed-loop iterative fusion of multi-scale features, enabling deep semantic information to be fed back to high-resolution features multiple times. This enhances the semantic response of small target regions and alleviates the problems of insufficient deep semantic transmission, scattered semantic response, and cross-scale semantic drift in existing single top-down feature fusion.

[0026] 3. Improve the ability to distinguish small targets in complex backgrounds.

[0027] This invention simultaneously utilizes frequency calibration to enhance target structural information and semantic calibration to enhance target semantic information, enabling the detection network to better distinguish between real small targets and background texture interference. For aerial images with small target sizes, complex backgrounds, dense target density, and significant occlusion, this invention can reduce missed detections and false detections, improving the reliability of detection results.

[0028] 4. Improve the stability of small target bounding box regression

[0029] This invention introduces gradient calibration loss in addition to the basic bounding box regression loss. By using the regression state scheduling factor and the center calibration potential function, the center correction gradient is enhanced when the predicted box has low overlap with the real box, no overlap, or a large center offset, so that the predicted box can get closer to the real target faster. When the positioning quality is improved, the influence of gradient calibration loss is weakened, and the boundary is refined by the basic regression loss, thereby improving the positioning accuracy of small targets.

[0030] 5. Alleviate performance deficiencies caused by the limited capacity of lightweight detector models

[0031] This invention does not rely on simply increasing network depth, expanding channel width, or introducing complex branches to improve detection accuracy. Instead, it addresses the problems of easily confused frequency details, diluted cross-scale semantics, and easily invalidated regression gradients in small target detection through targeted calibration. Therefore, this invention can improve small target detection performance while maintaining a low number of parameters and computational cost.

[0032] 6. Improve the model's deployment applicability on resource-constrained devices.

[0033] This invention employs a lightweight frequency calibration module, a semantic calibration neck network, and a plug-and-play gradient calibration loss, without significantly increasing the overall complexity of the detection network. It is suitable for scenarios with high requirements for storage, power consumption, and real-time performance, such as drones, vehicle-mounted devices, embedded devices, mobile terminals, and edge computing platforms.

[0034] 7. Improve the overall robustness of small target detection methods

[0035] This invention performs collaborative calibration of lightweight detectors from three levels: feature representation, feature fusion, and training optimization. This enables the model to maintain more stable detection performance under conditions such as complex backgrounds, scale changes, dense targets, occlusion, and noise interference, thereby improving the application value of small target detection methods in real-world scenarios. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a schematic diagram of the overall structure of the lightweight small target detection model provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the frequency calibration module structure provided in an embodiment of the present invention; Figure 3 This is a system block diagram of the electronic device provided in the embodiments of the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0039] First, it should be noted that in the embodiments of the present invention, the words "exemplarily," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplarily" is intended to present the concept in a specific manner. Furthermore, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either one or the other.

[0040] First Embodiment

[0041] To address the technical problems existing in the prior art, this embodiment provides a lightweight small target detection method based on calibration-driven methods, aiming to solve the following technical problems: First, it addresses the problem of unreliable frequency detail representation in lightweight small target detection, enabling the detection network to enhance edge, contour, and texture information related to small targets while suppressing high-frequency background interference.

[0042] Secondly, it addresses the problem of existing methods blindly enhancing high-frequency information, enabling input features to achieve target detail enhancement and background clutter suppression through frequency sub-band decomposition and adaptive calibration.

[0043] Third, it addresses the problem of insufficient cross-scale semantic fusion in small target detection, enabling deep semantic information to be fed back to high-resolution features multiple times, thereby alleviating the problems of semantic dilution, semantic drift, and target response dispersion.

[0044] Fourth, it addresses the problem of unstable gradient supervision during the regression of small target bounding boxes, enabling the predicted boxes to still obtain effective center correction gradients even when there is low overlap, no overlap, or large center offset.

[0045] Fifth, it addresses the challenge of balancing accuracy and efficiency in complex scenarios with lightweight detectors, improving the accuracy, stability, and deployment applicability of small target detection without relying on large-scale model expansion.

[0046] To address the aforementioned technical issues, this embodiment uses a lightweight target detection network as its foundation, introducing a frequency calibration module into the backbone network; a semantic calibration neck network into the neck network; and a gradient calibration loss during the training phase. These improvements aim to enhance frequency detail representation, cross-scale semantic fusion, and bounding box regression optimization processes in small target detection. This lightweight small target detection method can be applied to scenarios such as UAV aerial imagery, remote sensing imagery, intelligent transportation, security monitoring, robot perception, and edge device visual perception. The method can be implemented using an electronic device, which can be a terminal or a server. The execution flow of this method includes the following steps: S1, acquire the image to be detected; S2, the acquired image to be detected is input into a preset lightweight small target detection model; wherein, the lightweight small target detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract image features, the neck network is used to fuse features at different scales; the detection head is used to complete target classification, target confidence prediction, and bounding box regression; and the lightweight small target detection model introduces a calibration mechanism to achieve calibration of at least one direction of the feature representation, feature fusion, and bounding box regression gradient in small target detection; S3, output the detection results using the lightweight small target detection model.

[0047] Specifically, in this embodiment, the lightweight small target detection model is as follows: Figure 1As shown, it is based on a lightweight target detection network. This lightweight target detection network can be a lightweight detection network based on CSPDarknet, or a lightweight detection network combining a YOLO series lightweight detection network, a MobileNet series backbone network with a detection head, or other lightweight target detection networks including a backbone network, a neck network, and a detection head. In one specific embodiment, the lightweight target detection network removes the P5 detection branch with a 32x downsampling and adds a high-resolution P2 detection branch with a 4x downsampling, forming a multi-scale detection pyramid including P2, P3, and P4 to enhance small target detection capabilities.

[0048] A frequency calibration module is introduced into the backbone network, which uses discrete wavelet transform to decompose the input features into low-frequency and high-frequency sub-bands, and performs adaptive calibration and residual injection on different frequency components. A semantic calibration neck network is introduced into the neck network, which continuously calibrates multi-scale features through multi-round closed-loop semantic calibration and channel-level adaptive modulation fusion. Gradient calibration loss is introduced during the training phase, which dynamically adjusts the center correction intensity according to the overlap quality between the predicted box and the ground truth box. This improves the accuracy, localization stability and deployment applicability of small target detection in complex backgrounds while maintaining a lightweight structure. This improves the frequency detail representation, cross-scale semantic fusion and bounding box regression optimization process in small target detection.

[0049] Based on the above, the detection process using the lightweight small target detection model in this embodiment is as follows: Step 1: Multi-scale basic feature extraction; Acquire the image to be detected and input it into the backbone network of the lightweight object detection network.

[0050] The backbone network includes an initial feature extraction layer and multiple feature extraction stages. After convolution, normalization, activation, and downsampling, the image to be detected is processed to obtain basic feature maps at different scales.

[0051] The basic feature map includes high-resolution shallow features, intermediate-scale features, and low-resolution deep features. Among them, high-resolution shallow features are used to preserve the edge, contour, texture, and spatial location information of small targets; low-resolution deep features are used to provide semantic information related to the target category.

[0052] This step yields the multi-scale basic features required for subsequent frequency calibration and cross-scale semantic fusion.

[0053] Step 2: Characteristic subband decomposition based on the frequency calibration module; A frequency calibration module is set up in one or more feature extraction stages of the backbone network.

[0054] For the feature map of the input frequency calibration module, discrete wavelet transform is used to decompose it into subbands to obtain low-frequency subbands and multiple high-frequency subbands.

[0055] The low-frequency subband is used to represent smooth structures and main semantic components in the feature map; the high-frequency subband is used to represent horizontal edges, vertical edges, diagonal textures, target contours, and local fine-grained structures in the feature map.

[0056] This step decomposes the original spatial domain features into different frequency components, enabling subsequent modules to process semantic and detail information separately.

[0057] Step 3: Low-frequency semantic preservation and high-frequency detail calibration, such as Figure 2 As shown, it includes: For the low-frequency subband, a lightweight channel gating structure is used to generate low-frequency calibration weights, which are then used to adjust the low-frequency subband to maintain stable semantic information. Specifically, global average pooling is performed on the input feature map to obtain a channel description vector; this channel description vector is then sequentially input into a first-point convolutional or fully connected layer, a nonlinear activation function, a second-point convolutional or fully connected layer, and a sigmoid function to obtain low-frequency channel calibration weights with values ​​ranging from 0 to 1; these low-frequency channel calibration weights are then broadcast spatially and multiplied channel-by-channel with the low-frequency subband to obtain the calibrated low-frequency subband.

[0058] For high-frequency subbands, a lightweight convolutional thinning structure is first used to refine local features, and then high-frequency calibration weights are used to selectively enhance or suppress the high-frequency subbands. Specifically, based on the global context information of the input feature map, high-frequency channel calibration weights are generated through global average pooling, multilayer perceptron, nonlinear activation function, and sigmoid function. The high-frequency subbands first pass through a lightweight convolutional thinning structure consisting of depthwise convolution and pointwise convolution to obtain refined high-frequency features; then, the high-frequency channel calibration weights are broadcast in the spatial dimension and multiplied channel-by-channel with the refined high-frequency features to obtain the calibrated high-frequency subbands.

[0059] The lightweight channel gating structure can include a global average pooling layer, a multilayer perceptron, a nonlinear activation function, and a sigmoid function. The data flow of the lightweight channel gating structure is as follows: first, global average pooling is performed on the input feature map to obtain the global response of each channel; then, the global response is input into a first mapping layer for channel compression or transformation; next, it passes through a nonlinear activation function to obtain a nonlinear channel response; subsequently, it is input into a second mapping layer to restore the channel dimension to match the sub-band to be calibrated; finally, the channel calibration weights are obtained by normalization using the sigmoid function. After being broadcast spatially, the channel calibration weights are multiplied channel-by-channel with the corresponding frequency sub-band. The lightweight convolutional refinement structure can include depthwise convolution and pointwise convolution. In the lightweight convolutional refinement structure, multiple high-frequency sub-bands are concatenated to form high-frequency features. These high-frequency features are first input into a depthwise convolutional layer to perform local spatial refinement within each channel, suppressing disordered high-frequency noise; then, they pass through a nonlinear activation function and are input into a pointwise convolutional layer to fuse information between high-frequency sub-bands in different channels and directions, obtaining the refined high-frequency features.

[0060] This step enhances the edge, contour, and fine-grained texture information of small targets while suppressing high-frequency background interference such as road markings, building textures, shadows, vegetation textures, and sensor noise.

[0061] Step 4: Frequency calibration feature reconstruction and residual injection; The calibrated low-frequency and high-frequency subbands are input into the inverse wavelet transform to reconstruct the spatial domain frequency calibration residual signal.

[0062] The frequency calibration residual signal is fused with the input characteristics of the frequency calibration module to obtain the frequency-calibrated output characteristics.

[0063] In a specific implementation, a residual injection coefficient can be set to control the injection intensity of the frequency calibration residual signal.

[0064] Through this step, the present invention can enhance the structural details of small targets while avoiding the destruction of the overall semantic stability of the original features.

[0065] Step 5: Initial pyramid construction based on semantic calibration of the neck network; The multi-scale features output from the backbone network are input into the semantic calibration neck network.

[0066] First, channel alignment is performed on features at different scales to give each feature a preset channel dimension, forming an initial pyramid feature state.

[0067] The initial pyramid feature states include high-resolution feature states, intermediate-resolution feature states, and low-resolution semantic feature states.

[0068] This step provides a unified multi-scale feature representation for subsequent closed-loop cross-scale semantic calibration.

[0069] Step Six: Closed-loop cross-scale semantic calibration; In the semantic calibration neck network, multiple rounds of closed-loop semantic calibration are performed on the initial pyramid feature state.

[0070] In each round of calibration, the high-resolution features, intermediate-scale semantic features, and deep semantic features from the previous round are scale-aligned. This scale alignment can be achieved through upsampling, downsampling, or interpolation operations.

[0071] Adaptive fusion of scale-aligned multi-source features yields high-resolution calibrated features for the current round.

[0072] Subsequently, based on the high-resolution calibration features of the current round, the pyramid feature states at each scale are updated through upsampling, downsampling, and lateral connections, and the updated feature states are input into the next round of calibration.

[0073] This step transforms the traditional top-down semantic propagation in the feature pyramid into a closed-loop, multi-round iterative semantic calibration, enabling deep semantic information to be repeatedly fed back into high-resolution features. This alleviates the problems of semantic response dispersion and semantic dilution in small targets and enhances the semantic response of small target regions.

[0074] Step 7: Channel-level adaptive modulation fusion; In the closed-loop cross-scale semantic calibration process, channel-level fusion weights are generated based on the global context information of features at different scales.

[0075] The channel-level fusion weights are used to perform weighted aggregation of high-resolution features, intermediate-scale features, and deep semantic features.

[0076] When the background texture of the image is complex, the fusion weight of deep semantic features is increased to enhance the target category discrimination ability; when the location details of small targets are weak, the fusion weight of high-resolution features is increased to preserve the spatial positioning information of the target.

[0077] Through this step, the present invention can dynamically adjust the contribution of features at different scales according to the content of the input image, reducing semantic dilution and background interference problems caused by fixed fusion methods.

[0078] Step 8: Target classification and bounding box prediction based on the detection head; The multi-scale calibration features output from the semantic calibration neck network are input into the detection head.

[0079] The detection head performs target classification prediction, target confidence prediction, and bounding box regression prediction on feature maps of different scales.

[0080] For each candidate location, the detection head outputs the corresponding class confidence score, target presence probability, and predicted bounding box parameters. The predicted bounding box parameters can include the coordinates of the bounding box center point, width, and height, and can also include the coordinates of the top-left and bottom-right corners of the bounding box.

[0081] Candidate detection boxes are selected based on category confidence and target confidence, and duplicate detection boxes are removed by non-maximum suppression to obtain the final small target detection results.

[0082] Step 9: Bounding box regression optimization based on gradient calibration loss; During the training phase, gradient calibration loss is introduced on top of the basic bounding box regression loss.

[0083] The basic bounding box regression loss can be the cross-union loss, generalized cross-union loss, distance cross-union loss, full cross-union loss, or other bounding box regression losses.

[0084] The gradient calibration loss includes the regression state scheduling factor and the central calibration potential function.

[0085] In one specific implementation, let the predicted bounding box be... The true bounding box is ,in:

[0086]

[0087] in, Indicates the coordinates of the center point of the prediction box. Indicates the width and height of the prediction box; This represents the coordinates of the center point of the true bounding box. This represents the actual width and height of the bounding box.

[0088] The gradient calibration loss It is obtained by multiplying the regression state scheduling factor and the central calibration potential function, and its calculation formula is as follows:

[0089] in, Represents the prediction box With real frame The overlap quality index between the predicted and ground truth boxes can be IoU, GIoU, DIoU, CIoU, or other indexes that can characterize the overlap quality between the predicted and ground truth boxes. This represents the regression state scheduling factor, which is used to dynamically adjust the weight of the center correction term based on the overlap quality between the predicted box and the true box. This represents the center calibration potential function, which provides center correction constraints based on the offset distance between the center of the predicted box and the center of the true box.

[0090] In a preferred embodiment, IoU is used as the overlap quality index, and the regression state scheduling factor is:

[0091] in, Represents the prediction box With real frame The intersection-over-union ratio. When the overlap quality between the predicted bounding box and the ground truth bounding box is low. Smaller A larger value enhances the influence of the center correction term; when the overlap quality between the predicted box and the ground truth box is high... Larger The value of is reduced, thereby reducing the influence of the center correction term, so that the basic bounding box regression loss is mainly used for boundary refinement.

[0092] The central calibration potential function is:

[0093] in, This represents the center offset after normalization to the true bounding box size, and its calculation formula is:

[0094] in, Indicates the center point of the prediction box. Indicates the center point of the true bounding box. This represents the Euclidean distance between the center point of the predicted bounding box and the center point of the ground truth bounding box. This represents the scale normalization factor of the true frame.

[0095] Therefore, the specific form of the gradient calibration loss is as follows:

[0096] Furthermore, the lightweight small target detection model's bounding box regression optimization objective during the training phase is a weighted combination of the basic bounding box regression loss and the gradient calibration loss:

[0097] in, The basic bounding box regression loss can be IoU loss, GIoU loss, DIoU loss, CIoU loss, or other bounding box regression losses. The weight coefficients representing the gradient calibration loss are used to control the contribution of the gradient calibration loss to the overall bounding box regression optimization objective.

[0098] For multiple positive bounding box predictions within a training batch, the overall bounding box regression loss can be expressed as:

[0099] in, This indicates the number of positive samples participating in the bounding box regression training. Indicates the first One prediction box, Indicates the relationship with the first The predicted box matches the ground truth box.

[0100] Using the above calculation method, when the predicted bounding box and the ground truth bounding box are in a state of low overlap, no overlap, or large center offset, the regression state scheduling factor maintains a large effect of gradient calibration loss, thereby enhancing the correction gradient that moves the center of the predicted bounding box closer to the center of the ground truth bounding box. When the overlap quality between the predicted and ground truth bounding boxes improves, the regression state scheduling factor automatically reduces the influence of gradient calibration loss, allowing the basic bounding box regression loss to continue boundary refinement. Simultaneously, since the center offset is normalized using the ground truth bounding box scale, for small targets, the same pixel-level center offset will produce a larger normalized offset, thus obtaining stronger center correction constraints, which is beneficial for improving the stability and localization accuracy of small target bounding box regression.

[0101] Step 10: Model Training and Inference Deployment; During training, the training images are input into the calibration-driven lightweight detection network, which sequentially performs multi-scale basic feature extraction, frequency calibration, closed-loop semantic calibration, detection prediction, and loss calculation. The network parameters are then optimized based on classification loss, target confidence loss, basic bounding box regression loss, and gradient calibration loss.

[0102] After training, the resulting detection model will be deployed to drones, vehicle-mounted devices, embedded devices, mobile terminals, or edge computing platforms.

[0103] During the inference phase, the model takes the image to be detected as input, performs forward inference, and outputs the target category, bounding box location, and detection confidence.

[0104] In summary, the calibration-driven lightweight small target detection method in this embodiment achieves joint calibration of small target frequency details, cross-scale semantic information, and bounding box regression gradients without relying on large-scale model expansion through the synergistic effect of the frequency calibration module, semantic calibration neck network, and gradient calibration loss. This improves the accuracy, stability, and lightweight deployment applicability of small target detection in complex scenarios.

[0105] Second Embodiment

[0106] This embodiment provides an electronic device, such as... Figure 3 As shown, the electronic device includes a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment described above. Furthermore, the electronic device may also include a transceiver, the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0107] Below, in conjunction with Figure 3 A detailed introduction to each component of this electronic device is provided below: The processor is the control center of the electronic device. The electronic device may include multiple processors, each of which can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The term "processor" can refer to a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), other general-purpose processors, application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), one or more field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor can perform various functions of the electronic device by running or executing software programs stored in memory and by calling data stored in memory.

[0108] In a specific implementation, as one example, the processor may include one or more CPUs, for example... Figure 3 CPU0 and CPU1 shown are, of course, merely illustrative examples.

[0109] The memory is used to store the software program that executes the solution of the present invention, and the processor controls its execution. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.

[0110] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory may be integrated with the processor or exist independently, and may be accessed through the interface circuit of the electronic device ( Figure 3 (Not shown in the image) is coupled to the processor; however, this embodiment of the invention does not impose specific limitations on this.

[0111] The transceiver may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver can be integrated with the processor or exist independently, and can be connected through the interface circuit of the electronic device (…). Figure 3 (Not shown in the image) is coupled to the processor, and this embodiment of the invention does not specifically limit this.

[0112] In addition, it should be noted that, Figure 3 The structure of the electronic device shown is not intended to limit the device. Actual devices may include more or fewer components than shown, or combine certain components, or have different component arrangements. Furthermore, the technical effects achieved by this electronic device when performing the method of the first embodiment described above can be referenced to the technical effects described in the first embodiment; therefore, they will not be repeated here.

[0113] Third Embodiment

[0114] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc. The instruction stored therein can be loaded and executed by a processor in a terminal.

[0115] Furthermore, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention can take the form of a completely or partially hardware embodiment, a completely or partially software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented in software, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any usable medium accessible to a computer or a data storage device such as a server or data center containing one or more sets of usable media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive (SSD).

[0116] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0118] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element. Furthermore, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Additionally, the character " / " in this text generally indicates an "or" relationship between the preceding and following objects, but it can also indicate an "AND / OR" relationship. Please refer to the context for specific interpretations. "At least one" refers to one or more items, while "more than" refers to two or more items. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can be represented as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0119] Furthermore, it is understood that in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0120] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0121] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of functional modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Additionally, the functional units in the various embodiments of this invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0122] If the method is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0123] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention. It should be pointed out that although preferred embodiments of the present invention have been described, those skilled in the art, once they understand the basic inventive concept of the present invention, can make several improvements and modifications without departing from the principles described herein. These improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A lightweight small target detection method based on calibration-driven approach, characterized in that, include: Acquire the image to be detected; The acquired image to be detected is input into a preset lightweight small target detection model; wherein, the lightweight small target detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract image features, and the neck network is used to fuse features of different scales; The detection head is used to complete target classification, target confidence prediction, and bounding box regression; and the lightweight small target detection model introduces a calibration mechanism to calibrate at least one direction of the feature representation, feature fusion, and bounding box regression gradient in small target detection. The detection results are output using the lightweight small target detection model.

2. The lightweight small target detection method based on calibration-driven method as described in claim 1, characterized in that, The backbone network includes multiple feature extraction stages, and a frequency calibration module is set in one or more feature extraction stages. After the image to be detected is input into the backbone network, it undergoes convolution, normalization, activation and downsampling processing to obtain basic feature maps of different scales. In the feature extraction stage with the frequency calibration module, the basic feature maps are input into the frequency calibration module for frequency calibration to obtain frequency-calibrated output features. The frequency calibration module performs the following frequency calibration process: For the feature map of the input frequency calibration module, discrete wavelet transform is used to decompose it into subbands to obtain low-frequency subbands and multiple high-frequency subbands. For the low-frequency subband, a lightweight channel gating structure is used to generate low-frequency calibration weights, and the low-frequency subband is adjusted using the low-frequency calibration weights to obtain the calibrated low-frequency subband. For the high-frequency subband, a lightweight convolutional thinning structure is first used to refine the local features, and then high-frequency calibration weights are used to selectively enhance or suppress the high-frequency subband to obtain the calibrated high-frequency subband. The calibrated low-frequency subband and high-frequency subband are subjected to inverse wavelet transform to reconstruct the spatial domain frequency calibration residual signal. The spatial domain frequency calibration residual signal is residually fused with the feature map of the input frequency calibration module to obtain the frequency-calibrated output feature.

3. The lightweight small target detection method based on calibration-driven method as described in claim 2, characterized in that, The lightweight channel gating structure includes a global average pooling layer, a multilayer perceptron, a nonlinear activation function, and a sigmoid function.

4. The lightweight small target detection method based on calibration-driven method as described in claim 2, characterized in that, The lightweight convolutional refinement structure includes depthwise convolution and pointwise convolution.

5. The lightweight small target detection method based on calibration-driven method as described in claim 1, characterized in that, The neck network includes a closed-loop cross-scale semantic calibration mechanism. The process by which the neck network fuses features at different scales includes: The system receives multi-scale features output from the backbone network, performs channel alignment processing on features of different scales, and gives each scale feature a preset channel dimension to form an initial pyramid feature state. The initial pyramid feature state includes a high-resolution feature state, an intermediate-resolution feature state, and a low-resolution semantic feature state. Multiple rounds of closed-loop semantic calibration are performed on the initial pyramid feature state. In each round of calibration, the high-resolution feature state, intermediate-resolution feature state, and low-resolution semantic feature state from the previous round are first scale-aligned. The scale-aligned multi-source features are then adaptively fused to obtain the high-resolution calibration features for the current round. Subsequently, based on the high-resolution calibration features of the current round, the pyramid feature states at each scale are updated through upsampling, downsampling, and lateral connections. The updated feature states are then input into the next round of calibration. After multiple rounds of closed-loop semantic calibration, the fusion of features at different scales is achieved.

6. The lightweight small target detection method based on calibration-driven method as described in claim 5, characterized in that, The adaptive fusion of scale-aligned multi-source features includes: Channel-level fusion weights are generated based on global contextual information of features at different scales; among them, the more complex the image background texture, the greater the fusion weight of the low-resolution semantic feature state; the weaker the details of the small target location, the greater the fusion weight of the high-resolution feature state. Using the channel-level fusion weights, the scale-aligned high-resolution feature states, intermediate-resolution feature states, and low-resolution semantic feature states are weighted and aggregated.

7. The lightweight small target detection method based on calibration-driven method as described in claim 5, characterized in that, The scale alignment is achieved through upsampling, downsampling, or interpolation operations.

8. The lightweight small target detection method based on calibration-driven method as described in claim 1, characterized in that, The lightweight small target detection model optimizes the bounding box regression during the training phase by using a weighted combination of bounding box regression loss and gradient calibration loss; wherein the gradient calibration loss includes a regression state scheduling factor and a center calibration potential function. The regression state scheduling factor dynamically adjusts the weight of the center correction term based on the overlap quality between the predicted and ground truth boxes; the lower the overlap quality between the predicted and ground truth boxes, the greater the weight of the center correction term. The center calibration potential function generates a center offset penalty based on the distance between the center of the predicted box and the center of the ground truth box, and normalizes it according to the scale of the ground truth box, so that small targets can obtain a stronger center correction gradient under the same pixel offset.