Lightweight small target detection method based on machine vision

By introducing FasterNet module, VoV-GSCSP module, MPDIoU loss function and LAMP pruning algorithm in the YOLOv8 model, the problems of network complexity and calculation cost in resource-constrained environments of the YOLOv8 model are solved, and the efficient and accurate detection effect of small object detection is achieved.

CN120198772AInactive Publication Date: 2025-06-24山西省智慧交通实验室有限公司 +3
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510262307.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There is still room for improvement in the network complexity and computing cost of the YOLOv8 model in resource-constrained environments, and it is difficult to achieve feature extraction, data imbalance, and detection accuracy and speed balance in small object detection.

Method used

By introducing the FasterNet module instead of the C2f module in YOLOv8, partial convolution is used to reduce the computational amount and memory access; VoV-GSCSP module is introduced into the neck network, and multi-scale features are efficiently fused with GSConv and cross-level partial connection technology; bounding box regression performance is optimized using MPDIoU loss function based on the minimum point distance; adaptive pruning is used to reduce the model parameter amount and calculation amount.

Benefits of technology

The detection efficiency and accuracy of small object detection are significantly improved, the number of model parameters is reduced by about 50%, the calculation cost is reduced, and the bounding box regression performance is improved, making it suitable for real-time detection applications of resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005300202400000031
    Figure BDA0005300202400000031
  • Figure BDA0005300202400000034
    Figure BDA0005300202400000034
  • Figure BDA0005300202400000037
    Figure BDA0005300202400000037
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and deep learning, and particularly relates to a lightweight small target detection method based on machine vision, which comprises the following steps of: replacing a C2f module in YOLOv8 with a Faster Net module, and performing convolution operation on partial channels of an input feature map by utilizing partial convolution; a VoV-GSCSP module is introduced into the neck network, and efficient fusion is carried out on multi-scale features by combining GSConv and a cross-level part connection technology; by adopting a minimum point distance loss function, the regression performance of the bounding box is optimized by minimizing the distances between the prediction box and the left upper corner and the right lower corner of the real box; carrying out adaptive pruning according to the importance of the weight of each layer of the network by using an LAMP pruning algorithm; and training is carried out on the optimized network architecture. The mAP50 of the improved model on a TinyPerson data set reaches 46.3%, the recall rate is increased to 30%, and the detection performance is remarkably superior to that of a reference model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and deep learning, and particularly relates to a lightweight small target detection method based on machine vision. Background Art

[0002] Small target detection is a long-term challenge in computer vision, aiming to detect and locate small-sized and low-resolution targets in images. With the rapid development of drone technology, its applications in monitoring, detection and other fields are becoming increasingly widespread. In drone detection tasks, the targets are usually small and tiny targets in complex backgrounds, and traditional detection algorithms show the following limitations when dealing with small target detection: Difficult feature extraction: The resolution of small targets is low, and the available feature information is limited. Traditional feature extraction methods are difficult to extract strongly discriminative features, which affects the detection performance. Data imbalance problem: The number of samples of small targets is small in many actual scenarios, and it is easy to cause data imbalance during the training process, thus affecting the generalization ability of the model. Detection accuracy and speed balance: Complex network structures can improve detection accuracy, but have high requirements for computing resources and are not suitable for real-time application scenarios.

[0003] In the prior art, the YOLO series, as a representative of one-stage detection algorithms, has been widely used in real-time small target detection due to its efficient end-to-end target detection ability. However, although the YOLOv8 model has improved in terms of accuracy and speed, its network complexity and computational cost still have room for improvement in resource-constrained environments. In addition, how to further optimize the network structure to improve the detection ability for tiny targets remains a research hotspot. Summary of the Invention

[0004] Aiming at the technical problem that the network complexity and computational cost of the YOLOv8 model still have room for improvement in resource-constrained environments, the present invention provides a lightweight small target detection method based on machine vision, which combines lightweight design and optimized feature extraction, significantly improves the detection efficiency and accuracy, and provides new ideas and technical support for the field of small target detection.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0006] A lightweight small target detection method based on machine vision, comprising the following steps:

[0007] S1. By using the FasterNet module to replace the C2f module in YOLOv8, performing convolution operations on some channels of the input feature map using partial convolutions, reducing the amount of computation and memory access, and at the same time enhancing the diversity of the feature map and the feature representation ability;

[0008] S2. Introduce the VoV-GSCSP module into the neck network, combine the GSConv and cross-stage partial connection techniques to efficiently fuse multi-scale features, and at the same time reduce the number of feature map channels to reduce the computational cost;

[0009] S3. Adopt the minimum point distance loss function, optimize the bounding box regression performance by minimizing the distance between the predicted box and the upper left and lower right points of the ground truth box, and improve the detection accuracy of small targets;

[0010] S4. Use the LAMP pruning algorithm to adaptively prune according to the importance of the weights of each layer of the network, reducing the number of model parameters and computational volume while maintaining the detection performance;

[0011] S5. Train on the optimized network architecture, use the small target dedicated dataset to adjust the model parameters, and deploy on resource-constrained devices to achieve efficient real-time detection.

[0012] The method of performing convolution operations on some channels of the input feature map using partial convolution in S1 is as follows:

[0013] The FasterNet module performs convolution calculations on some channels of the input feature map through partial convolution technology, only calculating some channels to reduce the computational volume and memory access volume; at the same time, PConv makes full use of all channel information of the input data to enhance the feature representation ability;

[0014] The FasterNet module adopts a residual connection structure, combines partial convolution with two point convolutions to form an efficient feature extraction module.

[0015] The method of efficiently fusing multi-scale features by combining the GSConv and cross-stage partial connection techniques in S2 is as follows:

[0016] Combined with the GSConv technology, by integrating the advantages of standard convolution and depthwise separable convolution, and evenly distributing the feature information to each part of the feature map through the shuffle operation;

[0017] The VoV-GSCSP module also performs grouped convolution and fusion on features of different resolutions through cross-stage partial connection, effectively reducing the number of channels and computational complexity.

[0018] The method of optimizing the bounding box regression performance by minimizing the distance between the predicted box and the upper left and lower right points of the ground truth box in S3 is as follows:

[0019] The MPDIoU loss function based on the minimum point distance is:

[0020]

[0021] Among them, w and h represent the width and height of the input image. represents the sum of the squared distances between two points at the upper left corner. represents the sum of the squared distances between two points at the lower right corner; among them,

[0022]

[0023] Among them, represents the upper left corner point of the ground truth box. represents the upper left corner point of the predicted box.

[0024]

[0025] Among them, represents the lower right corner point of the ground truth box. represents the lower right corner point of the predicted box.

[0026] MPDIoU is defined as the bounding box loss function as follows:

[0027] L MPDIoU = 1 - MPDIoU

[0028] Among them, L MPDIoU represents the bounding box loss function.

[0029] The method of adaptively pruning according to the importance of the weights of each layer in S4 is as follows:

[0030] By evaluating the importance of the weights of each layer in the network, the pruning rate is adaptively adjusted layer by layer. First, according to the magnitude of the weights within the network layer, the weights are sorted to select important weights and unimportant weights, and according to the set pruning ratio p l the pruning threshold Γ l is determined:

[0031] Γ l = percentile(|W l |, 100 × p l )

[0032] Among them, percentile(·, p) represents the p-th percentile of the weight magnitude;

[0033] Then, the pruning ratio of each layer is adaptively adjusted, and the weights with less influence are pruned, thereby effectively reducing the redundant parameters of the network;

[0034] After the pruning is completed, the network is fine-tuned by retraining to restore the detection performance that may be reduced due to pruning.

[0035] The method of adjusting model parameters using a dedicated small target dataset in S5 is as follows: Integrate the improved backbone network, neck network, loss function, and pruning algorithm into the YOLOv8 framework to form a complete small target detection algorithm process.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] In the improved model of the present invention, the mAP50 reaches 46.3% on the TinyPerson dataset, and the recall rate is increased to 30%. The detection performance is significantly better than the benchmark model. Through the FasterNet module and LAMP pruning, the number of model parameters of the present invention is reduced by about 50%, and the computational cost is significantly reduced. The MPDIoU loss function of the present invention performs excellently in complex scenarios and non-overlapping target detection, improving the accuracy of bounding box regression. Through innovative network design, loss function optimization, and model slimming technology, the present invention achieves a balance between accuracy and efficiency in small target detection technology, providing an efficient and reliable solution for related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and those of ordinary skill in the art can also obtain other implementation drawings according to the provided drawings without creative efforts.

[0039] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship, or adjustment of the size should still fall within the scope covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.

[0040] Figure 1 It is a schematic flowchart of the method of the present invention;

[0041] Figure 2 It is a structural diagram of the FasterNet Block of the present invention;

[0042] Figure 3 It is a structural diagram of the GS BottleNeck of the present invention;

[0043] Figure 4 It is a structural diagram of the VoV-GSCSP module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. These descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0045] The following will further describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0046] A lightweight small target detection method based on machine vision in this embodiment, as Figure 1 shown, includes the following steps:

[0047] Step 1: When improving the backbone network of YOLOv8, as Figure 2 shown, introduce the FasterNet module to replace the original C2f module.

[0048] The FasterNet module performs convolution calculations on some channels of the input feature map through partial convolution (PConv) technology, and only calculates some channels to reduce the amount of calculation and memory access. At the same time, PConv makes full use of all channel information of the input data to enhance the feature representation ability.

[0049] The FasterNet module adopts a residual connection structure, combines partial convolution with two pointwise convolutions to form an efficient feature extraction module.

[0050] Through the residual connection, the network can learn key features more efficiently and effectively reduce the model complexity.

[0051] The improved backbone network significantly reduces the amount of calculation and the number of parameters, and can better adapt to the small target detection task, especially showing high detection performance in low-resolution and complex background scenarios.

[0052] Step 2: When improving the neck network, as Figure 4 shown, introduce the lightweight VoV-GSCSP module to optimize the feature fusion ability and reduce the calculation complexity.

[0053] As Figure 3As shown in the figure, by combining the GSConv (Group Shuffle Convolution) technology, the advantages of standard convolution and depthwise separable convolution are integrated, and through the shuffle operation, the feature information is evenly distributed to each part of the feature map.

[0054] This feature enables GSConv to have efficient feature extraction ability while maintaining a low computational cost. To further improve the network performance, the VoV-GSCSP module also performs grouped convolution and fusion on features of different resolutions through cross-stage partial connection (CSP), effectively reducing the number of channels and the computational complexity.

[0055] In addition, through the multi-level feature fusion strategy, the network can extract and combine features from different scales, enhancing the detection ability for small targets.

[0056] After the above improvements, the neck network significantly improves the detection accuracy and feature expression ability while reducing the computational resource requirements.

[0057] Step 3: To improve the accuracy of bounding box regression, the present invention proposes an MPDIoU (Minimum Point Distance IoU) loss function based on the minimum point distance. The formula is as follows:

[0058]

[0059] where w and h represent the width and height of the input image, represents the sum of the squared distances of the two upper left corner points, represents the sum of the squared distances of the two lower right corner points. Among them,

[0060]

[0061] where, represents the upper left corner point of the ground truth box, represents the upper left corner point of the predicted box.

[0062]

[0063] where, represents the lower right corner point of the ground truth box, represents the lower right corner point of the predicted box.

[0064] MPDIoU as the bounding box loss function is defined as follows:

[0065] L MPDIoU = 1 - MPDIoU.

[0066] The traditional IoU loss function has a problem of performance degradation when dealing with non-overlapping targets. MPDIoU significantly improves the geometric alignment performance of the bounding box by calculating the minimum distance between the top-left and bottom-right points of the predicted box and the ground truth box.

[0067] The MPDIoU loss function can not only adapt to non-overlapping target scenarios but also optimize the overlapping area between the detection box and the actual target, ensuring that the model can fit the target more accurately.

[0068] Compared with other geometric optimization methods, the calculation process of MPDIoU is more simplified, reducing the additional computational cost and significantly improving the bounding box regression performance of small target detection.

[0069] In the case of complex backgrounds or severe occlusions, MPDIoU shows higher robustness, which helps the model achieve stable object detection in various scenarios.

[0070] Step 4: To further reduce the model complexity and adapt to resource-constrained devices, the present invention introduces the LAMP (Layer Adaptive Magnitude-based Pruning) pruning algorithm.

[0071] The core idea of pruning is to adaptively adjust the pruning rate layer by layer by evaluating the importance of the weights in each layer of the network. First, the weights are sorted according to the magnitude of the weights within the network layer, and the important weights and unimportant weights are selected. Then, according to the set pruning ratio p l the pruning threshold Γ is determined l :

[0072] Γ l = percentile(|W l |, 100×p l )

[0073] where percentile(·, p) represents the p-th percentile of the weight magnitude.

[0074] Then, the pruning ratio of each layer is adaptively adjusted, and the weights with less influence are pruned, thereby effectively reducing the redundant parameters of the network.

[0075] After pruning, the network is retrained for fine-tuning to restore the detection performance that may be reduced due to pruning.

[0076] Compared with the global unified pruning method, the LAMP algorithm can better retain the weights that have a greater impact on the model performance, while significantly reducing the number of model parameters and the amount of computation.

[0077] The pruned model not only significantly improves the degree of lightweight but also greatly enhances the inference speed, making it suitable for real-time object detection on resource-constrained devices.

[0078] Step 5: After integrating the above optimization modules, integrate the improved backbone network, neck network, loss function, and pruning algorithm into the YOLOv8 framework to form a complete small object detection algorithm process.

[0079] In this process, first, the optimized backbone network extracts features from the input image, fully mining and capturing the features of small objects in the image; then, the optimized neck network fuses and enhances features of different scales to further improve the detection ability of small objects; finally, optimize the bounding box fitting based on the MPDIoU loss function and quickly complete the inference through the lightweight network after pruning.

[0080] The improved algorithm has significant advantages in detection accuracy improvement, calculation cost reduction, and model lightweight, and can achieve efficient and accurate detection of small objects in complex scenarios.

[0081] The above only elaborates in detail on the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A lightweight small target detection method based on machine vision, characterized in that: The following steps are involved: S1. By using the FasterNet module to replace the C2f module in YOLOv8, partial convolution is used to perform convolution operations on some channels of the input feature map, reducing the amount of calculation and memory access, while enhancing the diversity and feature representation capabilities of the feature map; S2, introduce the VoV-GSCSP module into the neck network, combine GSConv and cross-level partial connection technology to efficiently fuse multi-scale features, and reduce the number of feature map channels to reduce the computational cost; S3, using the minimum point distance loss function to optimize the bounding box regression performance and improve the detection accuracy of small targets by minimizing the distance between the predicted box and the upper left and lower right corners of the real box; S4, using the LAMP pruning algorithm, adaptively pruning according to the importance of the weights of each layer of the network, while reducing the number of model parameters and calculations while maintaining detection performance; S5. Train on the optimized network architecture, tune model parameters using a dedicated dataset for small targets, and deploy on resource-constrained devices for efficient real-time detection.

2. According to the lightweight small target detection method based on machine vision according to claim 1, it is characterized in that: The method of using partial convolution in S1 to perform convolution operation on some channels of the input feature map is: The FasterNet module performs convolution calculations on some channels of the input feature map through partial convolution technology, and only calculates some channels to reduce the amount of calculation and memory access; at the same time, PConv makes full use of all channel information of the input data to enhance the feature representation capability; The FasterNet module adopts a residual connection structure, combining partial convolution with two point convolutions to form an efficient feature extraction module.

3. The lightweight small target detection method based on machine vision according to claim 1, characterized in that: The method for efficiently fusing multi-scale features by combining GSConv and cross-level partial connection technology in S2 is as follows: Combined with GSConv technology, it integrates the advantages of standard convolution and depth-wise separable convolution, and evenly distributes feature information to each part of the feature map through shuffle operation; The VoV-GSCSP module also performs group convolution and fusion of features of different resolutions through cross-level partial connections, effectively reducing the number of channels and computational complexity.

4. The lightweight small target detection method based on machine vision according to claim 1, characterized in that: In S3, the method for optimizing the bounding box regression performance by minimizing the distance between the predicted box and the upper left corner and the lower right corner of the real box is: The MPDIoU loss function based on the minimum point distance is: Among them, w, h represent the width and height of the input image, d 2 1 means the sum of the squares of the distances between the two points in the upper left corner, represents the sum of the squares of the distances between the two points in the lower right corner; in, Indicates the upper left corner of the real box, Indicates the upper left corner of the prediction box; in, Indicates the lower right corner of the real box, Indicates the lower right corner of the prediction box; MPDIoU is used as the bounding box loss function and is defined as follows: L MPDIoU =1―MPDIoU Among them, L MPDIoU represents the bounding box loss function.

5. The lightweight small target detection method based on machine vision according to claim 1, characterized in that: The method of adaptive pruning according to the importance of weights of each layer of the network in S4 is: By evaluating the importance of weights in each layer of the network, the pruning rate is adaptively adjusted by layer. First, the weights are sorted according to the magnitude of the weights in the network layer, and important and unimportant weights are selected. Then, the pruning rate is adjusted according to the set pruning ratio p. l Determine the pruning threshold Γ l : Γ l =percentile(|W l |,100×p l ) Among them, percentile(·,p) represents the pth percentile of the weight magnitude; Then, the pruning ratio of each layer is adaptively adjusted to prune the weights with less impact, thereby effectively reducing the redundant parameters of the network; After pruning is completed, the network is retrained for fine-tuning to restore the detection performance that may have been reduced due to pruning.

6. The lightweight small target detection method based on machine vision according to claim 1, characterized in that: The method for adjusting the model parameters using the small target dedicated data set in S5 is: integrating the improved backbone network, neck network, loss function and pruning algorithm into the YOLOv8 framework to form a complete small target detection algorithm process.

Citation Information

Patent Citations

  • Old people falling detection method based on improved YOLOv8 model

    CN118692139A

  • Multi-scale automatic driving target detection method based on cross-space learning

    CN119181080A

  • Lightweight PCB defect detection method and device based on improved YOLOv8n and storage medium

    CN119559178A