Small target detection method based on superpixel mask and dynamic kernel

Through the object detection method based on superpixel mask and dynamic core, the problem of low accuracy and efficiency of small object detection in large field of view environments is solved, and more efficient detection results are achieved.

CN120374936APending Publication Date: 2025-07-25XIDIAN UNIV

Patent Information

Application Number
CN202510399101.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In a large field of view environment, small object detection has problems with low detection accuracy and efficiency, especially because the small object has low resolution, is susceptible to background noise and has sparse distribution, making it difficult for the prior art to effectively extract discriminant information and low computing efficiency.

Method used

The object detection method based on superpixel mask and dynamic core is adopted, and the initial image block segmentation is performed through the greedy slice module, and multi-scale feature fusion is performed in combination with the dynamic core module to generate a foreground area mask independent of the category, and the receptive field is dynamically adjusted to adapt to the context information of different scale targets.

Benefits of technology

It effectively improves detection accuracy and efficiency, avoids the calculation overhead of spatially related features being split and useless background information, and improves the accuracy and speed of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374936A_ABST
    Figure CN120374936A_ABST
Patent Text Reader

Abstract

The invention provides a small target detection method based on a superpixel mask and a dynamic kernel. The small target detection method comprises the following implementation steps: acquiring a training and testing sample set; constructing a target detection network model based on the superpixel mask and the dynamic kernel, and carrying out iterative training on the target detection network model; and obtaining a small target detection result. The greedy slicing module carries out initial image block segmentation according to the center point and the scale, the defect that space correlation features are segmented due to segmentation is avoided, the dynamic kernel module carries out multi-scale feature fusion on each image block, the receptive field can be dynamically adjusted, wide-area context information of targets of different scales can be flexibly adapted, and the image segmentation efficiency is improved. The detection precision is effectively improved; a superpixel mask generator generates a foreground region mask irrelevant to the category for each extracted feature image, and a dynamic kernel module performs multi-scale feature fusion on each foreground image block containing a target cluster obtained by a greedy slicing module; and the influence of processing useless background information on calculation overhead when the multi-scale features are obtained is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a target detection method based on superpixel masks and dynamic kernels. Background Art

[0002] In the target detection task in a large field of view environment, since small targets with a resolution less than 32×32 occupy only very few pixels and are usually distributed in clusters, due to limited original pixel information of small targets, key features are easily lost, resulting in a decline in positioning and classification accuracy; and the effective features of small targets are easily interfered by background noise and neighboring instances, and it is difficult to retain discriminative information during the feature extraction process, exacerbating the detection difficulty; at the same time, due to the sparse distribution of small targets, the model performs unnecessary feature extraction in large background areas, severely limiting the computational efficiency. The shooting distance of large field of view images has a large span, the distance difference between the target and the camera is significant, and the scale of the target to be detected varies greatly, resulting in difficulty for a single model to balance the influence of network depth on the detection of large and small targets, and it is easy to miss or misdetect when there are both large and small targets in the same scene.

[0003] The main methods for detecting small targets in a large field of view are methods based on feature enhancement and methods based on multi-scale feature fusion. Among them, the method based on multi-scale feature fusion has obvious advantages in cross-scale target detection. For example, a patent application with the application publication number CN114998688A and the name "A large field of view target detection method based on an improved YOLOv4 algorithm" discloses a lightweight target detection method based on multi-scale feature fusion. The invention trains an improved YOLOv4 network model including a high-precision feature extraction sub-network, an enhanced feature multi-scale fusion sub-network, and a lightweight target classification sub-network, and performs target detection on the large field of view image to be recognized through the optimized network model, improving the detection accuracy and speed. However, when dealing with dense small targets in a large field of view, since the spatial correlation features split into fragments due to the image being segmented by a 3×3 matrix in the feature extraction sub-network of the invention, it will affect the further improvement of detection accuracy, and due to the feature extraction of the background part, the computational cost is relatively large. Summary of the Invention

[0004] The object of the present invention is to overcome the defects existing in the above-mentioned prior art, and propose a small target detection method based on superpixel masks and dynamic kernels to solve the technical problems of low detection accuracy and efficiency existing in the prior art.

[0005] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0006] (1) Obtain a training sample set and a test sample set:

[0007] Preprocess the K large field-of-view images obtained, which include multiple target categories, and label the targets in the N preprocessed large field-of-view images. Then, form a training sample set with the N preprocessed large field-of-view images and their labels, and form a test sample set with the remaining K - N preprocessed large field-of-view images, where K > 300.

[0008] (2) Construct an object detection network model based on superpixel masks and dynamic kernels:

[0009] Construct an object detection network model that includes a cascaded feature extraction module, a superpixel mask generator, a greedy slicing module, a dynamic kernel module, and a detection module, and the output end of the superpixel mask generator is also connected to the input end of the detection module; among them, the dynamic kernel module includes a multi-scale depthwise separable convolution unit, a spatial pooling and attention mechanism unit, and a feature fusion unit.

[0010] (3) Iteratively train the object detection network model based on superpixel masks and dynamic kernels:

[0011] Iteratively train the object detection network model with the training sample set to obtain a trained object detection network model.

[0012] (4) Obtain the object detection results based on superpixel masks and dynamic kernels:

[0013] Use the test sample set as the input of the trained object detection network model for forward propagation to obtain the object detection results corresponding to the K - N test samples.

[0014] Compared with the prior art, the present invention has the following advantages:

[0015] 1. In the process of training the object detection network model and obtaining the object detection results, the greedy slicing module divides the initial image blocks according to the center point and scale, avoiding the defect that the spatially correlated features caused by the division are split. The dynamic kernel module then performs multi-scale feature fusion on each image block, which can dynamically adjust the receptive field and flexibly adapt to the wide-area context information of different-scale targets, effectively improving the detection accuracy.

[0016] 2. When obtaining the detection results of each sample, the superpixel mask generator generates a foreground region mask independent of the category for each extracted feature map, and the dynamic kernel module performs multi-scale feature fusion on each foreground image block containing the target cluster obtained by the greedy slicing module, avoiding the influence of the useless background information that cannot be distinguished from the foreground features when obtaining multi-scale features in the prior art on the computational cost, and effectively improving the detection efficiency. Description of the Drawings

[0017] Figure 1 This is the implementation flowchart of the present invention.

[0018] Figure 2 This is the structural schematic diagram of the object detection network model of the present invention. Detailed implementation manners

[0019] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Referring to Figure 1 , the present invention includes the following steps:

[0021] Step 1) Obtain a training sample set and a test sample set

[0022] Preprocess K large field-of-view RGB images including multiple target categories in the VisDrone and UAVDT datasets, label the targets in N preprocessed large field-of-view images, then form a training sample set from the N preprocessed large field-of-view images and their labels, and form a test sample set from the remaining K - N preprocessed large field-of-view images, where In this embodiment, K = 38327 and N = 23258;

[0023] Among them, the specific steps for preprocessing the obtained K large field-of-view RGB images including multiple target categories are as follows: perform sliding window cropping on each large field-of-view image. The resolution of the large field-of-view image is 4000×3000. Perform sliding window cropping on each image. The size of the cropped image is 1080×540, and zero-padding is performed on the edges of the cropped image to achieve data augmentation for this image, obtaining K preprocessed large field-of-view images.

[0024] Step 2) Construct an object detection network model based on superpixel masks and dynamic kernels, and its structure is as Figure 2 shown:

[0025] Construct an object detection network model including cascaded feature extraction modules, superpixel mask generators, greedy slicing modules, dynamic kernel modules, and detection modules, and the output end of the superpixel mask generator is also connected to the input end of the detection module; among them:

[0026] The superpixel mask generator module includes cascaded superpixel segmentation units, foreground probability calculation units, spatial probability propagation units, pixel-level probability refinement units, and post-processing units;

[0027] The greedy slicing module includes cascaded object center point localization and size estimation units and dynamic boundary optimization units. Among them, the object center point localization and size estimation unit includes a parallel 3×3 max pooling layer and a 9×9 average pooling layer; the dynamic boundary optimization unit includes a coordinate statistical position set and offset calculation;

[0028] The dynamic kernel module includes cascaded multi-scale depthwise separable convolution units, a spatial pooling unit, and a feature fusion unit. The multi-scale depthwise separable convolution units include a 5×5 convolution layer and a 7×7 convolution layer. The spatial pooling and attention mechanism unit includes parallel average pooling layer and max pooling layer, cascaded convolution layer and Sigmoid activation layer. The feature fusion unit includes a spatial mask weighting layer and a 1×1 convolution feature fusion layer.

[0029] Step 3) Iteratively train the object detection network model based on the superpixel mask and the dynamic kernel:

[0030] (3a) Initialize the iteration number as a, the maximum iteration number as A, A≥1000. The weight and bias parameters in the object detection network model for the a-th iteration are w a , b a , and let a = 0;

[0031] (3b) Use N training samples randomly selected from the training sample set as the input of the object detection network model for forward propagation to obtain the detection results of N images;

[0032] (3b1) Input the n-th image into the feature extraction module to obtain the feature map of the original image. The superpixel segmentation unit divides each feature map extracted by the feature extraction module into T superpixel regions. The foreground probability calculation unit calculates the initial foreground probability p raw of each superpixel region; the spatial probability propagation unit optimizes p raw ; the pixel-level probability refinement unit refines the optimized foreground probability p new ; the post-processing unit performs Gaussian filtering on the refined probability p to obtain the foreground mask M corresponding to each training sample;

[0033] Among them, the specific steps for generating a mask for each extracted feature map are as follows:

[0034] The superpixel segmentation unit first performs superpixel segmentation on each feature map, divides the image into T superpixel regions, and calculates the color histogram of each superpixel and the histogram of the background region raw :

[0035]

[0036] Among them, represents the color histogram of the current superpixel, Represents the background region histogram of the current superpixel, ∑ represents the summation operation, and max() is the operation of taking the maximum value;

[0037] The spatial probability propagation unit further optimizes the foreground probability distribution through spatial probability propagation. For each superpixel, the new optimized probability is calculated

[0038]

[0039] Among them, α controls the weights of the initial probability and the neighborhood probability, N(i) represents the neighborhood of superpixel i, and ω ij represents the weight between superpixels i and j;

[0040] The pixel-level probability refinement unit introduces the position weight ω pos , which reflects the relative position of the pixel in the superpixel. The pixels closer to the center have higher weights. The superpixel-level probability is refined to each pixel, and the position weight calculation formula is:

[0041]

[0042] Among them, x pixel represents the coordinates of the pixel, x center represents the coordinates of the superpixel center, and R sp represents the radius of the superpixel;

[0043] The foreground probability p of the pixel introducing position information can be modeled as:

[0044] p = p sp ·(0.6ω pos +0.4)·edge atten

[0045] Among them, p sp , ω pos respectively represent the basic foreground probability of the feature image pixel point, the position weight, and edge atten represents the foreground probability of the feature map edge pixel;

[0046] The post-processing unit uses a 3×3 Gaussian filter to smooth the foreground probability map to reduce noise, and obtains the foreground mask M corresponding to each training sample. The calculation formula is:

[0047] M = G * p

[0048] Among them, G represents Gaussian filtering, p represents the foreground probability introducing position information, and * represents the convolution operation;

[0049] The superpixel mask generator generates a foreground region mask independent of categories for each extracted feature map, avoiding the impact of the useless background information that cannot be distinguished from the foreground features in the existing technology when obtaining multi-scale features on the computational cost, and effectively improving the detection efficiency.

[0050] (3b2) The target center point positioning and size estimation unit performs center point positioning and scale estimation on the foreground region mask M, and performs initial image block segmentation according to the center point and scale; the dynamic boundary optimization unit optimizes the segmented initial image blocks to obtain foreground image blocks containing the target clusters.

[0051] Among them, the greedy slicing module performs image block segmentation on the generated foreground region mask M, and the implementation steps are as follows:

[0052] First, based on the target mask M predicted by the SP-Masker, a 3×3 max pooling operation is used to quickly locate the set of target center points where each element corresponds to a local extreme point on the feature map; synchronously execute a 9×9 average pooling operation to statistically calculate the activation pixel density and generate a set of target size estimations Prioritize selecting the target s with the largest size i The corresponding center coordinates Initialize a slice with a fixed size, whose width W p and height H p are determined according to a certain ratio of the original size of the feature map. The initial boundary coordinates are calculated through the center point offset:

[0053]

[0054] Among them, represents the current target center point, W p represents the slice width, and H p represents the slice height;

[0055] Then perform dynamic boundary optimization, and statistically calculate the set of activation positions within the slice By calculating the minimum coordinate offset and adjust the upper left corner coordinates of the slice to eliminate the invalid area, and obtain the foreground image block containing the target cluster;

[0056] The greedy slicing module performs initial image block segmentation according to the center point and scale, avoiding the defect that the spatially associated features are split during segmentation, and effectively improving the detection accuracy.

[0057] (3b3) The dynamic kernel module first performs a 5×5 convolution on the foreground image patch containing the target cluster to obtain a small receptive field feature map, then performs a 7×7 convolution to obtain a large receptive field feature map, and realizes channel mixing of the small and large receptive field feature maps through 1×1 convolution; performs channel-level average pooling and max pooling on the multi-scale features, generates a two-channel spatial attention map through convolution, and uses the Sigmoid activation to generate a spatial selection mask; finally, weights the multi-scale features based on the spatial mask and realizes feature fusion through 1×1 convolution to obtain the feature map after dynamic context modeling;

[0058] Among them, the dynamic kernel module performs multi-scale feature fusion on each segmented image patch, and the specific steps are as follows:

[0059] The multi-scale depthwise separable convolution unit constructs a large kernel convolution through an explicit decomposition method, decomposes it into a group of depthwise separable convolutions with increasing kernel sizes and dilation rates, and the receptive field RF calculation formula of the large kernel convolution is:

[0060] RF1=k1,RF i =d i (k i -1)+RF i-1

[0061] Among them, k represents the convolution kernel size, and d represents the dilation rate of the dilated convolution;

[0062] The spatial pooling unit performs average pooling and max pooling on the feature maps with different receptive field ranges after splicing respectively, extracts their spatial relationship information, and the calculation formula is:

[0063]

[0064] Among them, SFD avg and SFD max represent the average and max pooling spatial feature descriptors respectively, represents the average pooling operation, represents the max pooling operation;

[0065] The feature fusion unit splices the features after spatial pooling, and uses the convolutional layer Conv 2→2 (·) to convert the pooled features into a two-channel spatial attention map SAM:

[0066]

[0067] Among them, [·] represents the splicing operation, and Conv 2→2 (·) represents the two-channel conversion convolutional layer;

[0068] Then apply the Sigmoid function to activate it, obtaining the independent spatial selection mask SAM corresponding to the decomposed large kernel; apply the corresponding spatial selection masks to the feature maps with different receptive field sizes for weighting, and fuse them through a convolutional layer, finally obtaining the feature map Y after dynamic context modeling:

[0069]

[0070] Among them, represents the feature map of the i-th scale, SAM i represents the spatial selection mask corresponding to the feature map of the i-th scale, ∑ represents the summation operation, and Conv(·) represents the fusion operation;

[0071] The dynamic kernel module performs multi-scale feature fusion on each foreground image patch containing the target cluster obtained by the greedy slicing module, can dynamically adjust the receptive field, flexibly adapt to the wide-area context information of different-scale targets, and effectively improves the detection efficiency and accuracy.

[0072] (3b4) The detection module receives the feature map after dynamic context modeling and the foreground region mask. The foreground region mask guides the detection module to focus on the foreground region of the image. The detection module detects the feature map and obtains the detection results of N training samples.

[0073] (3c) Adopt the IoU loss function, and calculate the loss value of the network model through each object detection result and its corresponding ground truth label Then adopt the gradient descent method and calculate through the chain rule for the weight parameter w a and the bias parameter b a partial derivatives and and according to for w a 、b a to update, where:

[0074]

[0075] Among them, X n 、Y n respectively represent the region of the model prediction target and the region of the true target corresponding to the n-th training sample, |·| represents the absolute value operation, ∩ and ∪ respectively represent the intersection and union operations, ∑ represents the summation operation, w a ′、b a ′ respectively represent the update results of w a 、b a , α represents the learning rate, respectively represent for w a 、 for ba Partial derivative operation.

[0076] (3d) Determine whether a ≥ A holds. If so, obtain the trained object detection network model. Otherwise, set a = a + 1 and execute step (3b).

[0077] Step 4) Obtain the object detection result based on the superpixel mask and dynamic kernel:

[0078] Use the test sample set as the input of the trained object detection network model for forward propagation to obtain the object detection results corresponding to K - N test samples.

[0079] Next, combined with the simulation results, the technical effects of the present invention will be further described:

[0080] 1. Experimental conditions and content:

[0081] The hardware platform for the simulation experiment is: the processor is Intel Core i9 - 13900K, the GPU model is NVIDIA GeForce RTX 4090, and the memory size is 32G; the software platform for the simulation experiment is: Ubuntu 20.04 operating system.

[0082] A comparative simulation is carried out between the present invention and an existing small object detection method in terms of two indicators: the average precision AP for measuring the accuracy of the model and the calculation efficiency FPS. The results are shown in Table 1.

[0083] The average precision AP represents the average value of the areas under the precision - recall curves of all classes obtained at a step size of 0.05 within the IoU threshold range from 0.5 to 0.95. Its calculation formula is:

[0084]

[0085] where r i is the sorted recall rate point, and p interp (r i+1 ) is the interpolated precision corresponding to the recall rate, and ∑ represents the summation operation.

[0086] FPS is used to measure the calculation efficiency of the model. This indicator represents the number of image frames that the model can process per second.

[0087] 2. Analysis of experimental results:

[0088] Table 1 Simulation comparison results

[0089]

[0090] Referring to Table 1, compared with the prior art, in the VisDrone dataset and the UAVDT dataset, the average precision AP and the computational efficiency FPS index of the present invention have both improved, indicating that the detection accuracy has been effectively improved.

Claims

1. A small target detection method based on superpixel masks and dynamic kernels, characterized in that, It includes the following steps: (1) Obtain a training sample set and a test sample set: Preprocess the obtained K large field-of-view RGB images including multiple target categories, label the targets in the N preprocessed large field-of-view images, then form a training sample set with the N preprocessed large field-of-view images and their labels, and form a test sample set with the remaining K - N preprocessed large field-of-view images, where K > 300, (2) Construct an object detection network model based on superpixel masks and dynamic kernels: Construct an object detection network model including a cascaded feature extraction module, a superpixel mask generator, a greedy slicing module, a dynamic kernel module, and a detection module, and the output end of the superpixel mask generator is also connected to the input end of the detection module; among them, the dynamic kernel module includes a cascaded multi-scale separable convolution unit, a spatial pooling unit, and a feature fusion unit; (3) Iteratively train the object detection network model: Iteratively train the object detection network model through the training sample set to obtain a trained object detection network model; (4) Obtain small object detection results: Use the test sample set as the input of the trained object detection network model for forward propagation to obtain the object detection results corresponding to K-N test samples.

2. The method according to claim 1, wherein In step (1), the preprocessing of the obtained K large field-of-view images including multiple object classes is realized as follows: Perform sliding window cropping on each large field-of-view image and pad zeros to the edges of the cropped images to achieve data augmentation of the image, and obtain K preprocessed large field-of-view images.

3. The method according to claim 1, wherein The object detection network model described in step (2), where: The superpixel mask generator includes a cascaded superpixel segmentation unit, a foreground probability calculation unit, a spatial probability propagation unit, a pixel-level probability refinement unit, and a post-processing unit; The greedy slicing module includes a cascaded object center point localization and size estimation unit and a dynamic boundary optimization unit.

4. The method according to claim 1, wherein In step (3), the iterative training of the object detection network model is realized as follows: (3a) Initialize the number of iterations as a, the maximum number of iterations as A, where A ≥ 1000. The weight and bias parameters in the object detection network model for the a-th iteration are w a , b a , and set a = 0; (3b) The feature extraction module extracts features from each training sample; the superpixel mask generator generates masks for each extracted feature map; the greedy slicing module performs image block segmentation on the generated foreground region mask M; the dynamic kernel module performs multi-scale feature fusion on each segmented image block; The detection module performs sparse convolution operations on the foreground region mask M generated by the pixel mask generator and the multi-scale feature map generated by the dynamic kernel module to obtain the detection results of N training samples; (3c) The Intersection over Union (IoU) loss function is adopted, and the loss value of the network model is calculated through each object detection result and its corresponding ground truth label. Then, the gradient descent method is adopted, through updating the weights and bias parameters w a , b a to obtain the object detection network model for this iteration. (3d) Judge whether a≥A holds. If so, obtain the trained object detection network model. Otherwise, set a=a + 1 and execute step (3b).

5. The method according to claim 4, wherein In step (3b), the superpixel mask generator generates masks for each extracted feature map, and the implementation steps are as follows: The superpixel segmentation unit divides each feature map extracted by the feature extraction module into T superpixel regions; The foreground probability calculation unit calculates the initial foreground probability p of each superpixel region raw ; The spatial probability propagation unit optimizes p raw ; the pixel-level probability refinement unit refines the optimized foreground probability p new ; the post-processing unit performs Gaussian filtering on the refined probability p to obtain the foreground mask M corresponding to each training sample: M = G * p p = p sp ·(0.6ω pos + 0.4)·edge atten Among them, p sp , ω pos respectively represent the basic foreground probability and position weight of the feature image pixel, edge atten represents the foreground probability of the edge pixel of the feature map, G represents Gaussian filtering, and * represents the convolution operation.

6. The method according to claim 4, wherein In step (3b), the greedy slicing module performs image block segmentation on the generated foreground region mask M, and the implementation steps are as follows: The object center point localization and size estimation unit locates the center point and estimates the scale of the foreground region mask M, and performs initial image block segmentation according to the center point and scale; the dynamic boundary optimization unit optimizes the segmented initial image blocks to obtain foreground image blocks containing object clusters.

7. The method according to claim 4, characterized in that In step (3b), the dynamic kernel module performs multi-scale feature fusion on each segmented foreground image block, and the implementation steps are as follows: The multi-scale separable convolution unit performs feature extraction on each foreground image patch at I scales and then conducts channel concatenation; the spatial pooling unit performs average pooling and max pooling on the image patches after channel concatenation respectively, and conducts channel concatenation on the average pooling result and the max pooling result; The feature fusion unit fuses the multiple pooled features output by the spatial pooling unit to obtain the multi-scale feature map Y of each image patch: where \(I\geq2\), denotes the feature map of the \(i\)-th scale, \(i\in[1, I]\), SAM i denotes the corresponding spatial selection mask, \(\sum\) represents the summation operation, and Conv(·) represents the fusion operation.

8. The method according to claim 4, wherein The loss value described in step (3c) The calculation formula is: Among them, X n , Y n respectively represent the region of the model prediction target and the region of the true target corresponding to the nth training sample. |·| represents the absolute value operation, ∩ and ∪ respectively represent the intersection and union operations, and ∑ represents the summation operation.

9. The method according to claim 4, characterized in that, The update of the weight and bias parameters w a , b a in step (3c) is performed according to the following update formula: where, w a ′, b a ′ respectively represent the updated results of w a , b a , α represents the learning rate, respectively represent the partial derivative operations with respect to w a , and the partial derivative operation with respect to b a .

Citation Information

Patent Citations

  • Large-view-field target detection method based on YOLOv4 improved algorithm

    CN114998688A

Cited By

  • Multi-scale image data enhancement method and device, equipment and storage medium

    CN121121025A

  • A multi-scale image data enhancement method, apparatus, device, and storage medium

    CN121121025B

  • Power transmission channel image ground object intelligent segmentation method and device

    CN121259459A