Embedded device small target detection system and method based on learnable down-sampling
Patent Information
- Application Number
- CN202610979665.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-25
AI Technical Summary
它们增加了模型的参数量和计算复杂度,对于海思3516这类计算和内存资源高度受限的嵌入式平台并不友好
1、本发明通过在输入端引入任务导向的可学习下采样模块,并创新性地加入基于SSIM的特征增强损失,直接约束下采样层在小目标区域保留与Gabor参考特征相似的结构信息。该方法从根本上缓解了传统缩放导致的小目标特征丢失问题,在小目标数据集上取得了显著的精度提升。
Smart Images

Figure CN122821090A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and embedded artificial intelligence technology, specifically to a small target detection system and method for embedded devices based on learnable downsampling. Background Technology
[0002] With the popularization of deep learning technology, target detection algorithms, represented by the YOLO series, have been widely used in embedded devices in fields such as security monitoring and intelligent transportation due to their excellent balance between speed and accuracy. However, in actual deployments, a common contradiction severely restricts detection performance: the contradiction between high-resolution image sensors and limited computing resources.
[0003] Traditional detection algorithms, such as those using embedded chips like the HiSilicon Hi3516, integrate high-performance neural network processing units (NPUs), but their input resolution is typically limited (e.g., 640x640). Modern camera modules, however, often output images at 1920x1080 (Full HD) or even higher resolutions. Existing solutions typically employ fixed algorithms like bilinear interpolation to directly scale the high-resolution image to the model's input size. This crude downsampling method leads to severe blurring and loss of image details, particularly the texture and edge features of small targets, during scaling, directly causing a significant decrease in the detection model's accuracy, especially for small targets.
[0004] While the industry has proposed various methods to improve the YOLOv8 model structure to enhance small object detection capabilities—such as introducing attention mechanisms, improving the feature pyramid network, or replacing downsampling modules in the backbone network (e.g., using SPD-Conv to avoid information loss)—these improvements primarily focus on the model's internal structure. They increase the number of parameters and computational complexity, making them unsuitable for embedded platforms like the HiSilicon 3516, which are highly resource-constrained in terms of computation and memory. Furthermore, when deploying models on HiSilicon platforms, the complex model structure is more prone to accuracy loss during low-bit-width quantization (especially 8-bit quantization), further exacerbating the performance degradation problem.
[0005] Therefore, there is an urgent need for a preprocessing technique that can effectively preserve key information of small targets in high-resolution source images without significantly increasing the computational burden on edge devices, and is compatible with model quantization processes. Summary of the Invention
[0006] The first objective of this invention is to overcome the shortcomings of existing technologies that simply scaling high-resolution images leads to the loss of small target features. This invention provides a small target detection method for embedded devices based on learnable downsampling. This method introduces a trainable convolutional downsampling layer at the front end of the target detection network to achieve intelligent compressed sensing sampling jointly optimized with the detection task. This reduces image resolution while adaptively preserving feature information that is crucial for small target detection.
[0007] The second objective of this invention is to provide a system adapted to the above method, ensuring that the optimized model can run efficiently and with high precision on embedded NPU platforms such as HiSilicon 3516.
[0008] To achieve the first technical objective of this invention, the invention adopts the following technical solution.
[0009] A method for small target detection in embedded devices based on learnable downsampling includes the following steps: Step S1: Construct a target detection model that integrates a learnable downsampling module; A learnable downsampling module is inserted at the input of the YOLOv8 model before the backbone network. The learnable downsampling module contains two parallel paths: a main path and a noise suppression branch. The main path consists of a convolutional layer with a stride greater than 1, which is used to learnably downsample the high-resolution image to the size required by the model and expand the channel dimension; The noise suppression branch consists of depthwise separable convolution, global average pooling, and fully connected layers, generating channel gating with the same spatial size as the main path output. This is used to estimate local noise levels and dynamically suppress the characteristic responses of noise regions.
[0010] Step S2: End-to-end joint training and optimization; The constructed object detection model was trained using a high-resolution image dataset containing small targets and a composite loss function. When loading the training dataset, the scaling and padding preprocessing step is skipped, and the original image is directly fed into the object detection model for training. During backpropagation, the weights of the backbone network, detection head, and learnable downsampling module are updated synchronously based on the gradient calculated by the composite loss function. Compared to fixed downsampling, the method of this invention skips the scaling and padding preprocessing step and adopts a data-driven approach. This not only avoids the loss of small target features caused by traditional fixed downsampling and achieves end-to-end joint optimization of downsampling and detection tasks, but also maintains the consistency of data distribution during training and deployment, improving the detection accuracy of the quantized model. Furthermore, it allows the model to learn how to "trade" information when compressing resolution, theoretically preserving features sensitive to small targets to the greatest extent possible, thereby significantly improving the average accuracy (AP) for small targets.
[0011] Specifically, the composite loss function is expressed as: ; ; In the above formula, Represents the composite loss function; The original classification and regression loss functions of the YOLOv8 model; Enhance the loss for the newly added features; These are the weighting coefficients for the feature enhancement loss; This represents the number of small targets in the current image. For index variables; It is a structural similarity index; The feature map output by the learnable downsampling module; For the first The ground truth bounding box of a small target in the feature map Binary mask of the mapped region; For reference feature map; This indicates element-wise multiplication; in, and Based on SSIM calculation, the structural similarity between the small target region in the output feature map of the learnable downsampling module and the reference feature generated by the Gabor filter is calculated, thus forcing the learnable downsampling module to preserve the texture details of the small target.
[0012] Specifically, the step of synchronously updating the weights of the backbone network, the detection head, and the learnable downsampling module based on the gradient calculated according to the composite loss function includes: Detection head weight The update comes from the partial derivative of the total loss with respect to the detector head output, and the detector head weights. The update process is represented as follows: ; In the above formula, The updated detection head weights; Expanding on the chain rule , The predicted result output by the detection head; Base learning rate; Backbone network weights The updates come from the convergence of two paths: one is through backpropagation via the detection head, and the other is through indirect propagation via the feature enhancement loss through the downsampling layer, and the backbone network weights. The update process is represented as follows: ; In the above formula, The updated backbone network weights; The deep features output by the backbone network; ; Learnable downsampling module weights The update employs an adaptive learning rate and incorporates dual supervision from downstream detection and feature enhancement tasks, enabling the learning of downsampling module weights. The update process is represented as follows: ; In the above formula, The updated learnable downsampling module weights; This represents the gradient of the detection loss with respect to the output feature map of the learnable downsampling module; This represents the gradient of the feature enhancement loss with respect to the output feature map of the learnable downsampling module; For adaptive learning rate, , For modulation factor, The variance of the current gradient. , The sliding window size for calculating the gradient variance; Let be the gradient value at step t-k+1; This represents the mean of the gradient within the sliding window.
[0013] It should be noted that, in this invention, the gradient refers to the partial derivative of the composite loss function with respect to the weights, which is used to indicate the direction and magnitude of weight adjustment.
[0014] Step S3: Conversion, quantization, and deployment of the target detection model; Step S31: Conversion of the target detection model; After training, the complete object detection model, including the learnable downsampling module, is exported in ONNX format. Low-bit-width integer quantization calibration is performed using the model conversion tool corresponding to the embedded platform to adapt to the inference engine of the target neural network processing unit. Step S32: Quantization of the target detection model; The quantization process of the target detection model includes two parts: weight quantization and activation value quantization. For weight tensors Symmetric quantization maps the weight tensor to the integer range. Quantized weight tensor Represented as: ; In the above formula, Scaling factor ; This indicates rounding to the nearest integer. This indicates that the value will be truncated to... interval; During dequantization, integers are restored to floating-point approximations: ; In the above formula, These are approximate weight values after dequantization; For activation values, since their distribution varies with the input data, an entropy-based calibration method is used to determine the optimal scaling factor by minimizing the information loss before and after quantization. : ; In the above formula, The optimal scaling factor for the activation value; This indicates the search for the minimum value of the function. Values; The probability distribution of floating-point activation values; To use scaling factor Distribution of quantized integer activation values; KL divergence is used to measure the difference between two distributions; During the calibration process, a representative subset of images is selected from the training set for forward propagation, and statistical information of the activation values of each layer is collected to determine the optimal scaling factor for each layer. The learnable downsampling module is used as the first layer of the model for end-to-end quantization. The input of the learnable downsampling module is the original image distribution. The output is the feature map distribution. The input and output are communicated via learnable downsampling module parameters. Association, represented as: ; In the above formula, This represents the mapping function for the learnable downsampling module; Scaling factor during calibration and Collaborative optimization with subsequent layers minimizes the overall quantization error, as expressed as: ; In the above formula, The optimal scaling factor for the input layer. The optimal scaling factor for the output layer; The first in the model network layer; For the model network The probability distribution of the floating-point activation values of the layer; To use the first Layer scaling factor Distribution of quantized integer activation values; After quantization, convolution operation It can be approximated as integer operations: ; In the above formula, Represents the integer summation result of the convolution operation; For channel indexing; This represents the total number of channels in the convolution kernel; Indicates the first Integer weights for each channel; Indicates the first Integer input for each channel; Specifically, since the learnable downsampling module has already undergone task-oriented training, the feature map distribution of the learnable downsampling module... It is naturally adapted to subsequent detection networks, thus making it easier to maintain accuracy during quantization, mathematically manifested as a smaller KL divergence: ; In the above formula, This indicates the KL divergence before and after introducing a learnable downsampling module into end-to-end quantization; This represents the KL divergence before and after quantization when using conventional preprocessing scaling and model-separated quantization. This explains the underlying reason why the method of the present invention suffers less accuracy loss after quantization.
[0015] Step S33: Deployment of the target detection model; The video stream output from the camera is preprocessed and converted into RGB format, then directly input into the learnable downsampling module of the target detection model. The module outputs an optimized feature map and completes subsequent inference to generate detection results. The deployment formula for the quantized object detection model is expressed as: ; In the above formula, Output the detection results as floating-point numbers; This is the scaling factor for the output layer of the object detection model.
[0016] To achieve the second technical objective of this invention, the following technical solution is adopted.
[0017] An embedded device small target detection system based on learnable downsampling, used to implement the above-mentioned embedded device small target detection method, includes: The image acquisition module is used to acquire high-resolution raw images; A target detection model integrating a learnable downsampling module, wherein the learnable downsampling module has a built-in noise suppression branch and is jointly trained end-to-end, can adaptively reduce the interference of image noise on detection while optimizing downsampling of high-resolution images to the input size required by the target detection model; the target detection model is deployed on an embedded neural network processing unit to receive the high-resolution original image and output the detection result.
[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention introduces a task-oriented, learnable downsampling module at the input and innovatively incorporates a SSIM-based feature enhancement loss, directly constraining the downsampling layer to retain structural information similar to Gabor reference features in small target regions. This method fundamentally alleviates the problem of small target feature loss caused by traditional scaling, achieving a significant accuracy improvement on small target datasets.
[0019] 2. The learnable downsampling module introduced in this invention adds only a very small number of parameters and inference time, meeting the real-time requirements of embedded platforms. At the same time, the method of this invention proposes an adaptive learning rate strategy based on gradient variance for the learnable downsampling module, which accelerates model convergence. Furthermore, as the first layer of the model, the learnable downsampling module can be jointly quantized end-to-end with the main network, avoiding the accuracy loss in the preprocessing stage and simplifying the deployment process.
[0020] 3. The present invention can learn the lightweight noise estimation and suppression branch integrated in the downsampling module, which can dynamically adjust the feature response according to the local noise level of the image, effectively reducing the false alarm rate in complex scenarios such as low light and sensor noise, while maintaining sensitivity to small targets.
[0021] 4. The system of this invention is a general preprocessing enhancement scheme that does not rely on a specific backbone network structure. It can be flexibly combined with other object detection technologies such as attention mechanisms and feature pyramid optimization as a "plug and play" module, thereby further releasing the performance potential of the model. Attached Figure Description
[0022] To provide a more intuitive understanding of the technical implementation of this invention, the accompanying drawings involved in the embodiments of this invention are briefly described below. These drawings are used to assist in illustrating the implementation methods and are not intended to limit the invention. Those skilled in the art can make derivative designs based on the drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the architecture of an embedded device small target detection system based on learnable downsampling according to the present invention; Figure 2 This is a flowchart illustrating a small target detection method for embedded devices based on learnable downsampling according to the present invention. Figure 3 This is an end-to-end flowchart of the addition of a learnable downsampling module in an embodiment of the present invention; Figure 4 This is a schematic diagram of the conversion and deployment of the target detection model in an embodiment of the present invention; Figure 5 This is a schematic diagram of the reasoning process of the target detection model in this embodiment of the invention.
[0024] In the figure, the black box area represents the traditional scaling process, and the red box area represents the function introduced in this invention to remove the crude downsampling. Detailed Implementation
[0025] To facilitate understanding and implementation of the present invention by those skilled in the art, the various steps of the method proposed in this invention are described in detail below. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various modifications or alterations to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0026] Example 1 like Figure 2 As shown, this embodiment discloses a small target detection method for embedded devices based on learnable downsampling, including the following steps: Step S1: Construct a target detection model that integrates a learnable downsampling module; A learnable downsampling module is inserted at the input of the YOLOv8 model before the backbone network. The learnable downsampling module contains two parallel paths: a main path and a noise suppression branch. The main path consists of a convolutional layer with a stride greater than 1, which is used to learnably downsample the high-resolution image to the size required by the model and expand the channel dimension; The noise suppression branch consists of depthwise separable convolution, global average pooling, and fully connected layers, generating channel gating with the same spatial size as the main path output. This is used to estimate local noise levels and dynamically suppress the characteristic responses of noise regions.
[0027] Step S2: End-to-end joint training and optimization; like Figure 3 As shown, in this embodiment, a high-resolution image dataset containing small targets is used in conjunction with a composite loss function to train the constructed target detection model; In this embodiment, the scaling and padding preprocessing step is skipped when loading the training dataset, and the original image is directly fed into the target detection model for training; during backpropagation, the weights of the backbone network, the detection head, and the learnable downsampling module are updated synchronously based on the gradient calculated by the composite loss function. Compared to fixed downsampling, the method of this invention skips scaling and padding preprocessing and data-driven approaches. This not only avoids the loss of small target features caused by traditional fixed downsampling and achieves end-to-end joint optimization of downsampling and detection tasks, but also maintains the consistency of data distribution during training and deployment, improving the detection accuracy of the quantized model. Furthermore, it allows the model to learn how to "trade" information when compressing resolution, theoretically preserving features sensitive to small targets to the greatest extent possible, thereby significantly improving the average precision (AP) for small targets.
[0028] Specifically, the composite loss function is expressed as: ; ; In the above formula, Represents the composite loss function; The original classification and regression loss functions of the YOLOv8 model; Enhance the loss for the newly added features; The balance coefficient for feature enhancement loss; This represents the number of small targets in the current image. For index variables; It is a structural similarity index; The feature map output by the learnable downsampling module; For the first The ground truth bounding box of a small target in the feature map Binary mask of the mapped region; For reference feature map; This indicates element-wise multiplication; in, and Based on SSIM calculation, the structural similarity between the small target region in the output feature map of the learnable downsampling module and the reference feature generated by the Gabor filter is calculated, thus forcing the learnable downsampling module to preserve the texture details of the small target.
[0029] Specifically, the step of synchronously updating the weights of the backbone network, the detection head, and the learnable downsampling module based on the gradient calculated according to the composite loss function includes: Detection head weight The update comes from the partial derivative of the total loss with respect to the detector head output, and the detector head weights. The update process is represented as follows: ; In the above formula, The updated detection head weights; Expanding on the chain rule , The predicted result output by the detection head; Base learning rate; Backbone network weights The updates come from the convergence of two paths: one is through backpropagation via the detection head, and the other is through indirect propagation via the feature enhancement loss through the downsampling layer, and the backbone network weights. The update process is represented as follows: ; In the above formula, The updated backbone network weights; The deep features output by the backbone network; ; Learnable downsampling module weights The update employs an adaptive learning rate and incorporates dual supervision from downstream detection and feature enhancement tasks, enabling the learning of downsampling module weights. The update process is represented as follows: ; In the above formula, The updated learnable downsampling module weights; This represents the gradient of the detection loss with respect to the output feature map of the learnable downsampling module; This represents the gradient of the feature enhancement loss with respect to the output feature map of the learnable downsampling module; For adaptive learning rate, , For modulation factor, The variance of the current gradient. , The sliding window size for calculating the gradient variance; Let be the gradient value at step t-k+1; This represents the mean of the gradient within the sliding window.
[0030] It should be noted that, in this invention, the gradient refers to the partial derivative of the composite loss function with respect to the weights, which is used to indicate the direction and magnitude of weight adjustment.
[0031] Step S3: Conversion and quantization of the target detection model; Step S31: Conversion of the target detection model; like Figure 4After training, the complete target detection model, including the learnable downsampling module, is exported in ONNX format. Low-bit-width integer quantization (8-bit integer quantization is used as an example in this embodiment) is calibrated using the model conversion tool corresponding to the embedded platform to adapt to the inference engine of the target neural network processing unit. Step S32: Quantization of the target detection model; The quantization process of the target detection model includes two parts: weight quantization and activation value quantization. For weight tensors Symmetric quantization maps the weight tensor to a range of signed 8-bit integers. Quantized weight tensor Represented as: ; In the above formula, Scaling factor ; This indicates rounding to the nearest integer. This indicates that the value will be truncated to... interval; During dequantization, integers are restored to floating-point approximations: ; In the above formula, These are approximate weight values after dequantization; For activation values, since their distribution varies with the input data, an entropy-based calibration method is used to determine the optimal scaling factor by minimizing the information loss before and after quantization. : ; In the above formula, The optimal scaling factor for the activation value; This indicates the search for the minimum value of the function. Values; The probability distribution of floating-point activation values; To use scaling factor Distribution of quantized integer activation values; KL divergence is used to measure the difference between two distributions; During the calibration process, a representative subset of images is selected from the training set for forward propagation, and statistical information of the activation values of each layer is collected to determine the optimal scaling factor for each layer. The learnable downsampling module is used as the first layer of the model for end-to-end quantization. The input of the learnable downsampling module is the original image distribution. The output is the feature map distribution. The input and output are communicated via learnable downsampling module parameters. Association, represented as: ; In the above formula, This represents the mapping function for the learnable downsampling module; Scaling factor during calibration and Collaborative optimization with subsequent layers minimizes the overall quantization error, as expressed as: ; In the above formula, The optimal scaling factor for the input layer. The optimal scaling factor for the output layer; The first in the model network layer; For the model network The probability distribution of the floating-point activation values of the layer; To use the first Layer scaling factor Distribution of quantized integer activation values; After quantization, convolution operation It can be approximated as integer operations: ; In the above formula, Represents the 32-bit integer summation result of the convolution operation; For channel indexing; This represents the total number of channels in the convolution kernel; Indicates the first 8-bit integer weights for each channel; Indicates the first Each channel has an 8-bit integer input; Based on the sum of the obtained integers The output is restored to floating-point output by the output layer scaling factor.
[0032] Specifically, since the learnable downsampling module has already undergone task-oriented training, the feature map distribution of the learnable downsampling module... It is naturally adapted to subsequent detection networks, thus making it easier to maintain accuracy during quantization, mathematically manifested as a smaller KL divergence: ; In the above formula, This indicates the KL divergence before and after introducing a learnable downsampling module into end-to-end quantization; This represents the KL divergence before and after quantization when using conventional preprocessing scaling and model-separated quantization. This explains the underlying reason why the method of the present invention suffers less accuracy loss after quantization.
[0033] Step S4: Deployment of the target detection model; like Figure 5 As shown, the video stream output by the camera is preprocessed and converted into RGB format, then directly input into the learnable downsampling module of the target detection model, which outputs the optimized feature map and completes subsequent inference to generate the detection result. The deployment formula for the quantized object detection model is expressed as: ; In the above formula, Output the detection result as a 32-bit floating-point number; This is the scaling factor for the output layer of the object detection model.
[0034] Example 2 Based on Example 1, this example provides a system for implementing the embedded device small target detection method based on learnable downsampling described in Example 1. The system includes: The image acquisition module is used to acquire high-resolution raw images; A target detection model integrating a learnable downsampling module, wherein the learnable downsampling module has a built-in noise suppression branch and is jointly trained end-to-end, can adaptively reduce the interference of image noise on detection while optimizing downsampling of high-resolution images to the input size required by the target detection model; the target detection model is deployed on an embedded neural network processing unit to receive the high-resolution original image and output the detection result.
[0035] Specifically, to achieve optimized downsampling from high-resolution images to the model input size, this embodiment inserts a learnable downsampling module before the backbone network of the YOLOv8 model; the structure and parameters of this module are defined through the model configuration file, and a specific modification example is as follows: # Example of modifying the YOLOv8 model configuration file (model.yaml) # Note: This module does not require any scaling or padding preprocessing; it directly receives the original image (any size). # Output a feature map of fixed size 640×640×64, compatible with subsequent backbone networks. backbone: # Learnable downsampling module: Input any size -> Output 640x640x64 - [-1, 1, LearnableDownsampleWithNoiseSuppression, [64, 5, 3]] # Subsequent standard YOLOv8 backbone network - [-1, 1, Conv, [128, 3, 2]] # 640 -> 320 - [-1, 3, C2f, [128, True]] - [-1, 1, Conv, [256, 3, 2]] # 320 -> 160 - [-1, 6, C2f, [256, True]] # ... The subsequent network structure remains unchanged. """ Learnable downsampling module (integrated noise suppression branch) Note: This module completely skips the preprocessing step and directly receives the original image (any size). After downsampling via learnable convolution, adaptive pooling is used to force a fixed output size of 640x640. Ensure compatibility with the YOLOv8 backbone network to achieve true end-to-end joint optimization.
[0036] """ import torch import torch.nn as nn class LearnableDownsampleWithNoiseSuppression(nn.Module): """ Learnable downsampling module, including: - Main path: Standard convolutional downsampling - Noise estimation branch: Depthwise separable convolution + channel attention, generating noise gating """ def __init__(self, in_channels=3, out_channels=64, kernel_size=5, stride=3): super().__init__() # Mirror padding preserves boundary information self.pad = nn.ReflectionPad2d(kernel_size / / 2) # Main path: Learnable downsampling convolution self.conv_main = nn.Conv2d( in_channels, out_channels, kernel_size=kernel_size, stride=stride, padding=0, bias=False ) # Noise estimation branch (lightweight) # Depthwise separable convolution (stride consistent with main path, spatially aligned) self.dw_conv = nn.Conv2d( in_channels, in_channels, kernel_size=3, stride=stride, padding=1, groups=in_channels, bias=False ) # Channel Attention: Global Average Pooling + Fully Connected Generation Gating self.gap = nn.AdaptiveAvgPool2d(1) self.fc = nn.Sequential( nn.Linear(in_channels, in_channels / / 4, bias=False), nn.ReLU(inplace=True), nn.Linear(in_channels / / 4, out_channels, bias=False), nn.Sigmoid() # Outputs a gate value between 0 and 1 ) # Batch normalization and activation (for main path output) self.bn = nn.BatchNorm2d(out_channels) self.act = nn.SiLU(inplace=True) # Adaptive pooling: Forces a fixed output size of 640x640 to ensure compatibility with the backbone network. self.adaptive_pool = nn.AdaptiveAvgPool2d((640, 640)) def forward(self, x): # Input: [B, 3, H, W], any size, no preprocessing required x_pad = self.pad(x) # Main path downsampling main_feat = self.conv_main(x_pad) # [B, C, H', W'] # Noise estimation branch noise_feat = self.dw_conv(x) # [B, 3, H', W'] noise_desc = self.gap(noise_feat).flatten(1) # [B, 3] gate = self.fc(noise_desc).unsqueeze(-1).unsqueeze(-1) # [B, C, 1,1] # Gated modulation: Low gate value in high-noise regions suppresses characteristic responses. suppressed_feat = main_feat * gate # Follow-up processing out = self.bn(suppressed_feat) out = self.act(out) # Force output fixed size 640x640 out = self.adaptive_pool(out) return out.
[0037] The technical effects of the method of the present invention will be further illustrated below through two examples.
[0038] Example 1: Small target detection in traffic scenes based on HiSilicon Hi3516DV300; This example addresses the need for small target detection of vehicles and pedestrians in intelligent traffic monitoring scenarios. It designs and implements a small target detection system based on the HiSilicon Hi3516DV300 embedded platform.
[0039] System Design and Model Training The system hardware uses a HiSilicon Hi3516DV300 development board, equipped with a full HD camera, and deployed at a traffic intersection to collect 1920×1080 resolution video streams in real time. To improve small object detection capabilities, referring to the modifications to the UltralyticsYOLOv8 official code in Example 2, a learnable convolutional downsampling layer is added to the front end of the backbone network in the model.yaml configuration file. The configuration parameters are [-1, 1, Conv, [64, 5, 3]], indicating 3 input channels, 64 output channels, a 5×5 kernel size, and a stride of 3. This structure allows the input high-resolution image (1920×1080×3) to be downsampled to a 640×640×64 feature map after passing through this layer, while simultaneously achieving spatial dimension compression and feature channel expansion through convolution operations.
[0040] The network was trained using publicly available datasets containing small targets such as vehicles and pedestrians, including UA-DETRAC and VisDrone, with image resolutions of 1920×1080. The datasets cover scenes with varying lighting, weather conditions, and traffic density. A modified configuration was used to initialize the model, loading pre-trained COCO weights (skipping mismatched weights in the first layer). During training, the weights of this layer were optimized end-to-end along with the entire detection network using the Adam optimizer, trained for 300 epochs on 8 GPUs. This enabled the network to actively preserve feature information beneficial for small target detection in traffic scenes, especially the texture and contour details of pedestrians, distant vehicles, and traffic signs, even as the resolution decreases.
[0041] Model conversion and deployment After model training, the data is exported as `model_with_learnable_downsample.onnx`. Using the HiSilicon ATC tool, the input nodes are specified as the input of the learnable downsampled module (1920×1080×3), and 16-bit and 8-bit quantization calibrations are performed. The calibration set uses a representative portion of the training set to ensure minimal accuracy loss in the quantized model. During quantization, the learnable downsampled module, as the first layer of the model, has its input-output range co-optimized with subsequent network layers, avoiding the mismatch between traditional preprocessing and the main quantization of the model. Finally, a model file adapted to the Hi3516DV300 NPU inference engine is generated.
[0042] During deployment, the camera sensor is configured to output YUV data at 1920×1080@30fps. In the application layer, the acquired YUV images are converted to RGB format and then directly fed into the learnable downsampling module in the NPU model. This achieves optimized compression from high resolution to the model input size, ultimately outputting detection bounding boxes and confidence scores for vehicles and pedestrians. The system is integrated into the HiSilicon media processing platform, forming a complete embedded intelligent vision solution.
[0043] Compared to traditional bilinear interpolation scaling or fixed-step downsampling, the theoretical advantages of the method in small target detection in traffic scenes are reflected in the following three aspects: Feature Preservation Capability: Traditional scaling methods uniformly resample all pixels. Small targets (such as distant pedestrians or vehicles) occupy only a few to a dozen pixels, and their edges and textures are easily smoothed and blurred during interpolation. This invention employs a learnable downsampling module, which adaptively enhances the response weights of small target regions through task-driven learning of convolution kernel parameters, ensuring that distinguishable structural information is still preserved in the downsampled feature map. Combined with SSIM-based feature enhancement loss, the module is explicitly constrained to mimic the ideal texture reference generated by the Gabor filter, thereby mathematically maximizing the structural similarity of small target regions before and after downsampling.
[0044] Noise Robustness: Traffic scenarios often encounter interference such as low illumination, rain, fog, and sensor noise. Traditional downsampling methods cannot distinguish between signal and noise, causing noise components to be compressed into the feature map, increasing the risk of false alarms. This invention integrates a lightweight noise suppression branch into the learnable downsampling module. It estimates local noise levels through depthwise separable convolutions and generates channel gating to dynamically suppress feature responses in high-noise regions. This branch adds only a minimal amount of computation but effectively reduces the false detection rate in complex environments while maintaining sensitivity to small targets.
[0045] Quantization friendliness: Embedded platforms typically require models to perform low-bit-width integer quantization. In traditional schemes, preprocessing scaling (such as bilinear interpolation) is located outside the model, and its output distribution cannot be jointly optimized with the quantization parameters of the first layer of the model, easily leading to large quantization errors. This invention incorporates a learnable downsampling module as the first layer of the model to participate in end-to-end quantization calibration. Its input-output range, together with subsequent network layers, determines the optimal scaling factor, minimizing the overall quantization error. Theoretical analysis shows that the cumulative sum of KL divergence in this invention is significantly lower than that of traditional separable quantization schemes, thus achieving higher detection accuracy in practical deployments.
[0046] Example 2: Small target (chicken) detection in poultry farming scenarios based on HiSilicon Hi3516CV610; This example demonstrates the design and implementation of a small target detection system based on the HiSilicon Hi3516CV610 embedded platform, designed to meet the flock monitoring needs in modern poultry farming scenarios.
[0047] System Design and Model Training The system hardware uses a Hi3516CV610 development board, equipped with a full HD camera, and deployed inside a chicken coop to collect 1920×1080 resolution video streams in real time. To improve small object detection capabilities, this example inserts a learnable downsampling module at the front end of the YOLOv8n model. The convolutional downsampling layer of this learnable downsampling module adopts a design with a convolution kernel size of 3×3 and a stride of 3, with 3 input channels (corresponding to RGB images) and 32 output channels. Through appropriate boundary padding, this structure allows the input high-resolution image (1920×1080×3) to be downsampled to a 640×640×32 feature map after passing through this layer. At the same time, spatial dimension compression and feature channel expansion are achieved through convolution operations.
[0048] During training, the weights of the learnable downsampling module are optimized end-to-end along with the entire detection network, enabling it to actively retain beneficial feature information for chicken detection, especially the texture and contour details of small target chickens, even as the resolution decreases. End-to-end training is performed using a self-built chicken image dataset that covers different lighting, poses, and occlusion conditions, and data augmentation is used to improve the model's robustness.
[0049] Model conversion and deployment After model training, the model is exported in ONNX format and calibrated using 8-bit integer quantization with the HiSilicon ATC tool to ensure compatibility with the Hi3516CV610's NNIE inference engine. During deployment, the YUV video stream output from the camera is preprocessed by VPSS and converted to RGB format, then directly input into the learnable downsampling module in the model. This achieves optimized compression from high resolution to the model input size, ultimately outputting the chicken's detection bounding box and confidence score.
[0050] Compared to traditional downsampling methods and existing model improvement schemes (such as replacing the first layer with SPD-Conv), the theoretical advantages of the method in poultry farming scenarios are reflected in the following aspects: Adaptability to densely packed scenes with small targets: In chicken coop images, chickens are often close together, and traditional downsampling can easily lead to feature overlap between adjacent individuals, resulting in missed detections. In this invention, the learnable downsampling module uses convolution with a stride greater than 1 for spatial dimensionality reduction. Its receptive field covers the local neighborhood, and through training, it can learn the boundary response patterns that distinguish adjacent individuals. Simultaneously, the feature enhancement loss directly acts on the mask of the small target region, forcing the module to maintain the contrast between individuals in the feature map, thereby theoretically improving the detection recall rate in dense scenes.
[0051] Balancing computational efficiency and accuracy: Compared to internal improvements in models like SPD-Conv, SPD-Conv preserves all pixel information through a space-to-depth transformation, avoiding information loss but significantly increasing the computational cost of subsequent networks (the number of feature map channels expands to four times). This invention uses only a lightweight, learnable downsampling module at the input (with approximately 1 / 3 to 1 / 2 the number of parameters of a conventional convolutional layer), while the subsequent backbone network maintains its original structure. Therefore, the incremental computational cost of this invention is far less than that of the SPD-Conv scheme, while achieving comparable or even better feature preservation results through task-oriented learning. This makes this invention more suitable for embedded platforms with extremely limited computational resources.
[0052] Practical Value of the Noise Suppression Branch: The aquaculture environment contains complex noises such as feed dust, water mist, and uneven lighting. Traditional downsampling cannot distinguish between noise and targets, causing noise features to be passed to subsequent networks, increasing false alarms. In this invention, the noise suppression branch estimates local noise levels through depthwise separable convolutions and generates channel gating to suppress feature responses in high-noise regions. This branch requires no additional supervision signals and is entirely driven by detection loss and feature enhancement loss, enabling it to learn which regions are "untrustworthy." Theoretical analysis shows that this branch can effectively reduce the false alarm rate, especially in high-noise frames.
[0053] Consistency in Quantization Deployment: Poultry farming scenarios often require continuous 24-hour monitoring, making model inference power consumption and stability crucial. This invention quantizes the learnable downsampling module along with the main network, avoiding the mismatch between preprocessing steps and model quantization parameters in traditional solutions. The output distribution of the quantized model is closer to that of the floating-point model, thus maintaining high detection accuracy and better numerical stability in practical deployments.
[0054] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for small target detection in embedded devices based on learnable downsampling, characterized in that, Includes the following steps: Step S1: Construct an object detection model with an integrated learnable downsampling module: Insert a learnable downsampling module into the input of the YOLOv8 model and before the backbone network. The learnable downsampling module contains two parallel paths: the main path and the noise suppression branch. Step S2, End-to-end joint training and optimization: The constructed object detection model is trained using a high-resolution image dataset containing small targets and a composite loss function; The scaling and padding preprocessing step is skipped when loading the training dataset, and the original image is directly fed into the object detection model for training. During backpropagation, the weights of the backbone network, the detection head, and the learnable downsampling module are updated synchronously based on the gradient calculated by the composite loss function. Step S3: Conversion, quantization, and deployment of the target detection model; Step S31: Conversion of the target detection model; After training, the complete object detection model, including the learnable downsampling module, is exported in ONNX format and low-bit-width integer quantization calibration is performed using the model conversion tool corresponding to the embedded platform. Step S32: Quantization of the target detection model; The quantization process of the object detection model includes two parts: weight quantization and activation value quantization. After quantization, the convolution operation can be approximated as integer operation, resulting in an integer summation result, which is then restored to a floating-point output through the output layer scaling factor. Step S4: Deployment of the target detection model; The video stream output from the camera is preprocessed and converted to RGB format. It is then directly input into the learnable downsampling module of the object detection model, which outputs an optimized feature map and completes subsequent inference to generate detection results.
2. The method for small target detection in embedded devices based on learnable downsampling according to claim 1, characterized in that, The composite loss function used in the training process in step S2 is expressed as follows: ; ; In the above formula, Represents the composite loss function; The original classification and regression loss functions of the YOLOv8 model; Enhance the loss for the newly added features; This is the balance coefficient for the feature enhancement loss, with a value range of [0.1, 0.5], used to adjust the relative importance of the detection loss and the feature enhancement loss; This represents the number of small targets in the current image. For index variables; It is a structural similarity index; The feature map output by the learnable downsampling module; For the first The ground truth bounding box of a small target in the feature map Binary mask of the mapped region; For reference feature map; This indicates element-wise multiplication.
3. The method for small target detection in embedded devices based on learnable downsampling according to claim 1, characterized in that, Step S2, which involves synchronously updating the weights of the backbone network, the detector head, and the learnable downsampling module based on the gradient calculated using the composite loss function, includes: Detection head weight The update comes from the partial derivative of the total loss with respect to the detector head output, and the detector head weights. The update process is represented as follows: ; In the above formula, The updated detection head weights; Expanding on the chain rule , The predicted result output by the detection head; Base learning rate; Backbone network weights The updates come from the convergence of two paths: one is through backpropagation via the detection head, and the other is through indirect propagation via the feature enhancement loss through the downsampling layer, and the backbone network weights. The update process is represented as follows: ; In the above formula, The updated backbone network weights; The deep features output by the backbone network; ; Learnable downsampling module weights The update employs an adaptive learning rate and incorporates dual supervision from downstream detection and feature enhancement tasks, enabling the learning of downsampling module weights. The update process is represented as follows: ; In the above formula, The updated learnable downsampling module weights; This represents the gradient of the detection loss with respect to the output feature map of the learnable downsampling module; This represents the gradient of the feature enhancement loss with respect to the output feature map of the learnable downsampling module; For adaptive learning rate, , For modulation factor, The variance of the current gradient. , The sliding window size for calculating the gradient variance; Let be the gradient value at step t-k+1; This represents the mean of the gradient within the sliding window.
4. The method for small target detection in embedded devices based on learnable downsampling according to claim 1, characterized in that, The weight quantization process described in step S32 is as follows: For weight tensors Symmetric quantization maps the weight tensor to the integer range. Quantized weight tensor Represented as: ; In the above formula, Scaling factor ; This indicates rounding to the nearest integer. This indicates that the value will be truncated to... interval; During dequantization, integers are restored to floating-point approximations: ; In the above formula, This is the approximate weight value after dequantization.
5. The method for small target detection in embedded devices based on learnable downsampling according to claim 1, characterized in that, The activation value quantization process in step S32 is as follows: For activation values, since their distribution varies with the input data, an entropy-based calibration method is used to determine the optimal scaling factor by minimizing the information loss before and after quantization. : ; In the above formula, The optimal scaling factor for the activation value; This indicates the search for the minimum value of the function. Values; The probability distribution of floating-point activation values; To use scaling factor Distribution of quantized integer activation values; KL divergence is used to measure the difference between two distributions; During the calibration process, a representative subset of images is selected from the training set for forward propagation, and statistical information of the activation values of each layer is collected to determine the optimal scaling factor for each layer. The learnable downsampling module is used as the first layer of the model for end-to-end quantization. The input of the learnable downsampling module is the original image distribution. The output is the feature map distribution. The input and output are communicated via learnable downsampling module parameters. Association, represented as: ; In the above formula, This represents the mapping function for the learnable downsampling module; Scaling factor during calibration and Collaborative optimization with subsequent layers minimizes the overall quantization error, as expressed in: ; In the above formula, The optimal scaling factor for the input layer. The optimal scaling factor for the output layer; The first in the model network layer; For the model network The probability distribution of the floating-point activation values of the layer; To use the first Layer scaling factor Distribution of quantized integer activation values.
6. The method for small target detection in embedded devices based on learnable downsampling according to claim 1, characterized in that, The deployment formula for the quantized target detection model in step S33 is expressed as follows: ; In the above formula, Output the detection results as floating-point numbers; This is the scaling factor for the output layer of the object detection model.
7. A small target detection system for embedded devices based on learnable downsampling, used to implement the method as described in any one of claims 1-6, characterized in that, include: The image acquisition module is used to acquire high-resolution raw images; A target detection model integrating a learnable downsampling module, wherein the learnable downsampling module has a built-in noise suppression branch and is jointly trained end-to-end, which can adaptively reduce the interference of image noise on detection while optimizing downsampling of high-resolution images to the input size required by the target detection model; The target detection model is deployed on an embedded neural network processing unit to receive the high-resolution original image and output the detection results.