A fast detection and robust tracking method for small moving targets
By replacing the backbone network with MobileNetV3 in the YOLOv5 network, adding a CBAM attention mechanism, adding a small object detection layer and optimizing the loss function, combined with the DeepSORT target tracking algorithm, the problems of insufficient accuracy, generalization ability and robustness in small object detection are solved, and efficient and accurate small object detection and tracking are achieved.
Patent Information
- Application Number
- CN202410889846.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-07-04
AI Technical Summary
The prior art has problems of insufficient detection accuracy, model generalization ability and robustness in small object detection, especially in the recognition and tracking of small objects in images or videos.
Improvements in the YOLOv5 network model include replacing the backbone network as MobileNetV3, adding a CBAM attention mechanism, adding a small object detection layer, and optimizing the loss function to Focal Loss and CIoU interpolation function, and using the DeepSORT target tracking algorithm.
It realizes the improvement of detection accuracy, enhancement of detection capabilities for small targets, reduce model parameters and calculation complexity in small target detection tasks, and improves the robustness and generalization capabilities of the model.
Smart Images

Figure CN119418085B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of target detection, and in particular relates to a method for rapid detection and robust tracking of weak moving targets. Background Art
[0002] Object detection is an important research direction in the field of computer vision and the basis of other complex visual tasks. Small object detection has long been a difficult point in object detection. It aims to accurately detect small objects with few visual features in images. Intelligent detection and tracking of small moving objects is an important research topic in the field of computer vision. In real scenes, small objects exist in large numbers, so small object detection has broad application prospects. As a subfield of general object detection, small object detection plays an important role in many fields such as autonomous driving, smart medical care, defect detection and aerial image analysis. In many scenarios, such as autonomous driving, drone navigation, video surveillance, etc., accurate detection and tracking of small moving objects is a key task. Small object detection has long been a difficult point and research hotspot in computer vision. Driven by deep learning, small object detection has made major breakthroughs and has been successfully applied to national defense security, intelligent transportation and industrial automation.
[0003] Nowadays, driven by deep learning, although significant progress has been made in the field of general object detection, the research progress of small object detection is relatively slow, which is one of the most challenging tasks in computer vision. Specifically, even with advanced detectors, there is still a huge gap in the performance of detecting small objects and regular-sized objects. The characteristics of small moving objects make them more difficult to identify and track in images or videos. Compared with regular-sized objects, small targets have small coverage, low resolution, less feature information, and insufficient feature expression in image or video data. Small targets usually lack sufficient appearance information, making it difficult to distinguish them from background or similar targets. Because the intrinsic structure of small targets leads to poor visual appearance and noisy representation, the feature representation of small objects is not ideal. The fundamental reason is their limited size and general feature extraction paradigm. In addition, the current popular feature extractors usually downsample the feature maps to reduce spatial redundancy and learn high-dimensional features, which will inevitably affect the detection of small targets.
[0004] The most important aspect of small target detection research is the accuracy of detection. Researchers are committed to finding ways to improve the accuracy of detection. Liu Siyuan et al. proposed a small target detection model with an attention mechanism based on YOLOv5l in the paper "Small Target Detection for Unmanned AerialVehicle Images Based on YOLOv5l". The paper proposed a module called C3ECA, which combines ECA with BottleneckCSP to help the new backbone network find areas with a large number of targets. In the patent "A Small Target Detection System and Method Based on Improved YOLOv5" (patent number: CN202311401134.X) published in 2023, Xi Yuanyuan et al. proposed adding a multi-scale purification module to the network to suppress conflicting information after multi-scale fusion. At the same time, a coordinate attention mechanism and a small target prediction head were added, which effectively enhances the detection accuracy, robustness and generalization ability of small targets. These methods have significantly improved detection capabilities, but their disadvantage is that they require relatively high computing resources because they involve a large number of parameters and calculations, have higher requirements for hardware equipment and deployment environment, and there is still much room for improvement in speed.
[0005] Nowadays, the detection of small targets requires fast and high accuracy, which is a challenging problem. This requires the use of efficient algorithms and models. Improving the detection accuracy without affecting the detection speed of small targets is the main research content of current researchers. Gao Tianyu et al. proposed a backbone network attention method in the document "Small Object Detection Methodbased on Improved YOLOv5", which enhances the characteristics of target features. And due to the lightweight characteristics of CBAM, it has little effect on the detection speed while improving the ability of the target to be accurately detected. In the patent "A Method and System for Small Target Detection of Improved YOLOv5 Network" published by Xiong Jingwen in 2021 (Patent No.: CN202111051345.6), it is proposed to add SE network units to the end of the main feature extraction network of the yolov5 network to form an improved yolov5 network model, and the attention mechanism is more lightweight. The advantage of making the model more lightweight is that while improving the accuracy of the original network, the speed is not affected. The disadvantage is that the generalization ability and robustness have not been greatly improved, and there is also a lot of room for improvement in detection accuracy. However, these methods need to be further improved in terms of speed, detection accuracy, and model generalization ability.
[0006] Object tracking is an important task in the field of computer vision, involving a variety of related background technologies. The development of object tracking technology has made significant progress, allowing computers to track objects in real time and accurately in videos. Correlation filter technology is one of the earliest and simplest methods in object tracking, but it has certain limitations when dealing with object deformation, occlusion, etc. The object tracking method based on deep learning uses convolutional neural networks to learn the feature representation of the object, which has strong adaptability and robustness. Multi-target tracking is an important research topic in object tracking, which aims to track multiple targets simultaneously and associate targets. Long-term object tracking is an important method to deal with situations such as long-term occlusion, disappearance and reappearance of the target. For example, a trajectory model can be used to predict the position of the target during occlusion. Online object tracking refers to object tracking without a pre-trained model, using only current observation data. This method usually uses an online update mechanism to adapt to changes in the appearance of the target.
[0007] At present, researchers have proposed many methods to address a series of problems with small and weak targets, such as data manipulation methods, scale perception methods, feature fusion methods, super-resolution methods, context modeling methods and other methods to deal with the difficulties of small targets. However, the detection of small and weak targets still has the above-mentioned problems such as insufficient detection accuracy, model generalization ability and robustness. Summary of the invention
[0008] In order to solve the above technical problems, the present invention proposes a method for rapid detection and robust tracking of weak moving targets to solve the problems existing in the above prior art.
[0009] To achieve the above object, the present invention provides a method for rapid detection and robust tracking of a small moving target, comprising:
[0010] A small target detection network is constructed based on the YOLOv5 network model, wherein the main structure in the small target detection network is the YOLOv5 network structure, wherein the backbone network structure and the prediction head structure of the YOLOv5 network structure are adjusted and a CBAM module is added, and the adjusted YOLOv5 network structure is trained to obtain a target detection network;
[0011] Acquire image and video data, detect the image and video data through a small target detection network to generate different target detection results, and track the target detection results through a tracking algorithm to generate a target tracking result.
[0012] Optionally, adjusting the backbone network structure includes:
[0013] The backbone network structure of the YOLOv5 network structure is replaced with the MobileNetV3 network structure, and the original convolution in the backbone network structure is replaced with a lightweight convolution;
[0014] The MobileNetV3 network structure includes an hourglass module, wherein the hourglass module includes two consecutive inverted residual modules connected in series to form an inverted residual structure, wherein the inverted residual structure is used to enhance the nonlinear capability of the model.
[0015] Optionally, adjusting the prediction head structure includes:
[0016] In the prediction head structure, a small target detection layer is added, wherein the backbone network outputs a 640*640 feature map, and the small target detection layer uses a 160*160 feature map to detect targets with pixels larger than 4*4.
[0017] Optionally, adjusting the backbone network structure and the prediction head structure further includes:
[0018] The CBAM module is added to the backbone network structure and the prediction head structure of the YOLOv5 network structure, wherein the CBAM module consists of a channel attention mechanism and a spatial attention mechanism.
[0019] Optionally, during the training of the adjusted YOLOv5 network structure, the classification loss function is replaced, wherein the replaced classification loss function adopts the Focal Loss loss function, and the Focal Loss loss function is:
[0020] FL(p t )=-α t (1-p t ) γ log(p t )
[0021] Among them, FL(p t ) represents the loss value of the Focal Loss loss function, α t Represents the weight coefficient related to the sample category, p t It represents the probability that the model predicts that the sample belongs to the correct category, and γ represents the adjustment parameter.
[0022] Optionally, during the training of the adjusted YOLOv5 network structure, the regression loss function is replaced, wherein the replaced regression loss function is a CIoU intersection-over-union function, and the CIoU intersection-over-union function is:
[0023]
[0024] Among them, L CIoUrepresents the loss value of the CIoU intersection-over-union function, IoU represents the ratio of the intersection and union of two given shapes, ρ represents the Euclidean distance between the center point of the predicted box and the true box, b and b gt They represent the coordinates of the center points of the predicted box and the real box respectively, c represents the diagonal length of the minimum circumscribed rectangle representing the two images, α represents the balance parameter, and v represents the aspect ratio.
[0025] Optionally, the tracking algorithm adopts the DeepSORT target tracking algorithm.
[0026] Compared with the prior art, the present invention has the following advantages and technical effects:
[0027] Based on the above improvements, the present invention can further optimize the yolov5 model to obtain more efficient and accurate small target detection results.
[0028] First, the backbone is replaced with MobileNetV3 to reduce the model parameters and computational complexity. MobileNetV3 is a lightweight network structure that extracts features by using hourglass blocks. The hourglass blocks use depthwise separable convolution and residual connections and perform feature reorganization in the channel direction, thereby reducing the amount of computation and parameters. In addition, the use of GhostConve can further reduce the number of model parameters. It reduces redundant calculations by sharing weights and improves the computational efficiency of the model.
[0029] The CBAM attention mechanism added next can improve detection accuracy. The CBAM attention mechanism introduces channel attention and spatial attention in the feature map. Channel attention strengthens important feature channels by learning channel weights, while spatial attention focuses on important spatial positions by learning spatial weights. This attention mechanism enables the model to pay more attention to important feature information, thereby improving detection accuracy.
[0030] The added small target detection layer enhances the detection capability of small targets. This prediction layer performs predictions at a shallower layer of the network, which can better capture the detailed features of small targets. By performing predictions at the small target prediction layer, the model can focus more on the detection task of small targets, improving the accuracy and recall rate of small targets.
[0031] Finally, the improved loss function CIoU takes into account the complete intersection, union, and diagonal distance between target boxes, and more accurately evaluates the prediction deviation of the target box. By using the CIoU loss function, the model can better optimize the prediction of the target box and further improve the accuracy of detection.
[0032] Based on the above improvements, the technical effects of the present invention include reducing the number of model parameters and computational complexity, improving detection accuracy, enhancing the detection capability of small targets, and improving the loss function. These technical effects enable Yolov5 to have higher accuracy, faster speed, and better robustness in small target detection tasks, so that the accuracy, generalization, and robustness of small target detection can be improved. This is of great significance for many practical application scenarios, such as unmanned driving, video surveillance, etc., and can provide users with better experience and services. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0034] Figure 1 Schematic diagram of a network residual structure according to an embodiment of the present invention;
[0035] Figure 2 is a target detection flow chart of an embodiment of the present invention;
[0036] Figure 3 This is a Yolov5+DeepSORT workflow diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0037] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0038] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0039] Based on the above-mentioned shortcomings of the existing technologies for detecting small and weak targets, such as insufficient detection accuracy, model generalization ability and robustness, the present invention proposes a lightweight multi-target detection and tracking algorithm based on YOLOv5+DeepSORT. The innovation of the present invention is that it is based on the YOLOv5-6.1 deep learning network, improves the original backbone part, adds a lightweight attention mechanism CBAM, adds a prediction head suitable for small target detection, optimizes the loss function, and finally optimizes the intersection-over-union function.
[0040] It should be noted that the present invention can serve in technical fields such as autonomous driving, drone navigation, and video surveillance, and its weak targets are targets whose horizontal or vertical pixels in the captured image are less than a certain ratio of the total pixels, and objects or people whose horizontal and vertical pixels are less than 16*16 pixels, including but not limited to people, vehicles, infrastructure, markers, small objects, etc. photographed from a distance. The image data or video data corresponding to the detection task can be determined by the corresponding detection task, and the improved network structure proposed by the present invention can be trained by the image data or video data and the corresponding labeled weak targets, and the trained network can be used to identify and track weak targets in unknown image data or video data.
[0041] like Figure 2 As shown, the detailed invention steps of the improved technical solution are as follows:
[0042] (1) Small Target Detection Model
[0043] The present invention constructs a data set through relevant preprocessing methods such as image standardization, data annotation and data enhancement, and constructs a target detection network based on the YOLOv5 network structure, wherein in the target detection network, the backbone network, i.e., the main network, uses GhostConve lightweight convolution, CBAM attention mechanism and MobileNetV3 with hourglass blocks for improvement, and in the neck network, uses GhostConve lightweight convolution and a multi-scale feature module with CBAM attention mechanism for improvement, in addition to the original large-scale detection layer, medium-scale detection layer and small-size detection layer in the detection layer, a micro-scale detection layer is added as a small target detection layer to detect small targets, the target detection network is trained through the data set, and the corresponding detection structure is output to generate a small target detection result.
[0044] 1) Improvements to the backbone of YOLOv5
[0045] The improvement of this part is first proposed by the present invention. Here, it is mainly to adapt to the detection of small targets and make the backbone network lightweight, so that the model is lightweight. First of all, small targets usually have a small size and low pixel density. Therefore, for small target detection tasks, the use of a lightweight backbone network structure can greatly reduce the number of model parameters and computational complexity, and improve the reasoning speed of the model. Secondly, this can also improve a certain degree of detection accuracy. The lightweight backbone network structure is specially designed and optimized to better adapt to the feature representation and detection requirements of small targets. Compared with the traditional backbone network structure, the lightweight backbone network structure can more effectively capture the detailed features of small targets and improve the detection accuracy and recall rate of small targets. Finally, this change can reduce the risk of overfitting. Small targets usually face the problem of data scarcity. By using a lightweight backbone network structure, the model can be better prevented from overfitting in small target detection tasks and the generalization ability of the model can be improved.
[0046] The details of the improvement are to replace the backbone part of YOLOv5-6.1 with MobileNetV3 with an hourglass (Sandglass) block, and replace all the original convolutions, that is, ordinary convolution layers in the replaced backbone part, MobileNetV3, with lightweight convolution GhostConve.
[0047] MobileNetV3 is a lightweight neural network with a low number of parameters and computational overhead. The MobileNetV3 network adopts a series of innovative designs, including inverted residual structure, separable convolution and linear bottleneck, to improve the performance and efficiency of the model. It can provide excellent accuracy and precision of target detection while maintaining a low computational cost. It is suitable for a variety of resource-constrained application scenarios. Experimental verification shows that this module greatly reduces the number of parameters and computational complexity, and the code running speed has been significantly improved.
[0048] like Figure 1 As shown in the figure, in order to retain enough useful information, MobileNetV3 proposed the Sandglass module. The Sandglass block consists of two consecutive inverted residual structures. The key idea of the inverted residual structure is to increase the nonlinear expression ability while keeping the model parameters small. The Sandglass block can extract features at a deeper level and increase the nonlinear ability of the model by connecting two inverted residual structures in series. This makes the model have stronger expression ability and better performance without changing its size. Figure 1 In which, the first structure is the residual structure, the second structure is the inverted residual structure, Pw is the point-by-point convolution, Dw is the separable convolution, Linear represents the linear layer, and ReLu represents the activation function.
[0049] This part also replaces the original convolutional layers in all backbones with the lightweight convolutional layer GhostConve. This operation can significantly reduce the number of model parameters and computational overhead, making the model more lightweight and increasing detection speed. In addition, GhostConv can also improve the generalization ability of the model, which helps to improve the performance of the model in various scenarios. Experimental verification shows that this operation can further improve the efficiency and applicability of the model without sacrificing performance. It can make the model more lightweight, efficient, and suitable for various resource-constrained application scenarios.
[0050] 2) Add a lightweight attention mechanism CBAM
[0051] Adding an attention mechanism to the network model of target detection can effectively improve the accuracy and robustness of the model. In order not to make the model complicated and affect the detection speed, a lightweight attention mechanism CBAM is added. This is a method based on spatial and channel attention mechanisms. It not only considers the relationship between channels, but also the relationship between spatial positions. In Yolov5, a CBAM block can be added after each convolutional layer to implement the attention mechanism. The CBAM block consists of a channel attention mechanism and a spatial attention mechanism.
[0052] Adding an attention mechanism is very suitable for solving the difficulties of small target detection. First, small targets usually have a small size and low pixel density, and are easily affected by image noise and background interference. Adding CBAM can enhance the feature representation ability of small targets and help the model better focus on the important features of the small target area. By adaptively adjusting channel attention and spatial attention, CBAM can effectively extract and strengthen the useful features of small targets. Secondly, the CBAM attention mechanism uses a combination of channel attention and spatial attention to better capture the local detail information of small targets. Channel attention can adaptively learn the correlation between different channels, while spatial attention can focus on the position of small targets, enabling the model to better perceive the detailed features of small targets. Then, by introducing the CBAM attention mechanism, the model's detection accuracy for small targets can be improved. CBAM can effectively enhance the discrimination and expression ability of small target features, enabling the model to more accurately classify and locate small targets, thereby improving the accuracy of small target detection. Finally, the CBAM attention mechanism is adaptive and can adjust the intensity of attention according to the size and complexity of the target. This means that CBAM can flexibly adapt to targets of various sizes, including small targets. Therefore, by introducing the CBAM attention mechanism, the model's detection performance for objects of different sizes can be improved.
[0053] Since the CBAM module is simple and efficient in design, it can be easily integrated into different layers of any CNN architecture, and the computational and parameter overheads brought by the introduction of the CBAM module are very small and almost negligible. While improving the performance of the model detection target, it does not affect the speed of the model. For the first time, the present invention adds CBAM to both the backbone network and the prediction head, and experimental verification shows that the detection accuracy is improved.
[0054] 3) Add a prediction head that is more suitable for small target detection (improve detection capability)
[0055] In order to detect small targets, the targets are small, numerous and densely packed with pixel sizes basically below 15*15. The downsampling multiple of YOLOv5 is relatively large. Therefore, learning feature information is a challenge for small objects in deep feature maps. So the algorithm is not suitable for small target detection. In the original YOLOv5 model, there were only three different detection layers. In the feature map with a training image scale of 640*640, the large detection layer of 80*80 can detect objects larger than 8*8 pixels, the medium detection layer with a 40*40 feature map can detect objects larger than 16*16 pixels, and the small detection layer with a 20x20 feature map can detect targets larger than 32*32 pixels. Add a small object detection layer for this situation. The feature map with a larger detection layer adds 160*160 to detect objects with pixels larger than 4*4.
[0056] This change is more targeted and meaningful for small target detection. First, because small targets usually have small size and low pixel density in images, traditional target detection algorithms may not be able to effectively detect these small targets. By adding a small detection head designed specifically for small targets, the model's detection performance for small targets can be improved, enabling it to detect and locate small targets more accurately. Secondly, the small detection head can optimize the features of small targets and improve the model's perception of small targets. By adjusting the network structure, loss function, and other aspects to specifically handle small targets, the model's sensitivity to small targets can be enhanced, thereby improving the accuracy of small target detection. Then, designing a small detection head for small targets can help reduce the model's false detection rate during the detection process and avoid misjudging the background or other irrelevant areas as small targets. The small detection head can analyze the features of small targets more carefully, thereby performing target detection more accurately and reducing the occurrence of false positives. Finally, by designing a separate small detection head for small targets, the complexity and computational complexity of the model can be reduced, and the speed of the training and reasoning process can be accelerated. The small detection head can handle the detection task of small targets more effectively and improve the efficiency and performance of the model.
[0057] Based on the combination of shallow feature maps and deep feature maps, small objects with pixels less than 15*15 have good detection effects.
[0058] 4) Loss function optimization
[0059] The classification loss function in the original code is replaced with the FocalLoss loss function. Focal Loss is a loss function for difficult-to-classify samples. It strengthens the learning of difficult-to-classify samples by increasing the weight of difficult-to-classify samples. Focal Loss is used as part of the classification loss function to replace the original cross entropy loss function. This loss function can help the model better learn difficult-to-classify samples, thereby improving the accuracy of the model. The expression of Focal Loss is as follows:
[0060] FL(p t )=-α t (1-p t ) γ log(p t ) (1)
[0061] Among them, FL(p t ) represents the loss value of the Focal Loss loss function, α t Represents the weight coefficient related to the sample category, p t It indicates the probability that the model predicts that the sample belongs to the correct category, and γ represents the adjustment parameter
[0062] 5) Optimization of intersection-union ratio function
[0063] The regression loss function in the original code was replaced with the CIoU intersection-over-union function, which is a further extension of GIoU. It takes into account the impact of the aspect ratio of the bounding box and more comprehensively measures the overlap between bounding boxes. In Yolov5, the CIoU loss function is used to replace part of the bounding box regression loss function to replace the original MSE loss function. The CIoU intersection-over-union function can better handle changes in the size and position of the bounding box, thereby improving the robustness and accuracy of the model. The CIoU intersection-over-union function is shown below:
[0064]
[0065] Among them, L CIoU represents the loss value of the CIoU intersection-over-union function, IoU represents the ratio of the intersection and union of two given shapes, ρ represents the Euclidean distance between the center point of the predicted box and the true box, b and b gt Represent the center point coordinates of the predicted box and the real box respectively, c represents the diagonal length of the minimum circumscribed rectangle representing the two images, α represents the balance parameter, and v represents the aspect ratio
[0066] Combining Focal Loss and CIoU loss functions can optimize classification and regression tasks at the same time and improve the overall performance of the model. This method enables the model to better learn difficult-to-classify objects while also more accurately predicting the location and size of the bounding box. However, it requires increased computation and adjustment of hyperparameters.
[0067] (2) Small Target Tracking Algorithm
[0068] For tracking small targets, the DeepSORT target tracking algorithm is used. This is a target tracking algorithm based on deep learning and trajectory sorting, which is widely used in multi-target tracking tasks in videos. It is an improvement on the classic SORT algorithm. DeepSORT combines deep learning technology with the SORT algorithm to achieve more accurate and robust multi-target tracking. Its main idea is to use the deep learning model to extract the feature representation of the target and track the target through trajectory matching and association. The overall design workflow is as follows Figure 3 shown.
[0069] In the process of detecting and tracking small targets, the present invention first obtains video data of small targets, performs frame-by-frame target detection on the video data, and in the target detection process, detects small targets in the video data by preparing a data set, improving the structure of the yolov5 network, and training the improved yolov5 network, i.e., the target detection network, outputs a detection frame, and tracks small and weak targets through the DeepSORT target tracking algorithm. First, Kalman filtering is used for prediction, and then motion matching (Mahalanobis space distance), appearance matching (Mahalanobis space distance), cascade matching, and IoU matching are performed on the prediction results and the detection frame in turn, and the tracker is updated, and the final target tracking result is output to achieve detection and tracking of small and weak targets.
[0070] Based on the above improvements, the present invention can further optimize the yolov5 model to obtain more efficient and accurate small target detection results.
[0071] First, the backbone is replaced with MobileNetV3 to reduce the model parameters and computational complexity. MobileNetV3 is a lightweight network structure that extracts features by using hourglass blocks. The hourglass blocks use depthwise separable convolution and residual connections and perform feature reorganization in the channel direction, thereby reducing the amount of computation and parameters. In addition, the use of GhostConve can further reduce the number of model parameters. It reduces redundant calculations by sharing weights and improves the computational efficiency of the model.
[0072] The CBAM attention mechanism added next can improve detection accuracy. The CBAM attention mechanism introduces channel attention and spatial attention in the feature map. Channel attention strengthens important feature channels by learning channel weights, while spatial attention focuses on important spatial positions by learning spatial weights. This attention mechanism enables the model to pay more attention to important feature information, thereby improving detection accuracy.
[0073] The added small target detection layer enhances the detection capability of small targets. This prediction layer performs predictions at a shallower layer of the network, which can better capture the detailed features of small targets. By performing predictions at the small target prediction layer, the model can focus more on the detection task of small targets, improving the accuracy and recall rate of small targets.
[0074] Finally, the improved loss function CIoU takes into account the complete intersection, union, and diagonal distance between target boxes, and more accurately evaluates the prediction deviation of the target box. By using the CIoU loss function, the model can better optimize the prediction of the target box and further improve the accuracy of detection.
[0075] Based on the above improvements, the technical effects of the present invention include reducing the number of model parameters and computational complexity, improving detection accuracy, enhancing the detection capability of small targets, and improving the loss function. These technical effects enable Yolov5 to have higher accuracy, faster speed, and better robustness in small target detection tasks. This is of great significance for many practical application scenarios, such as unmanned driving, video surveillance, etc., and can provide users with better experience and services.
[0076] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for rapid detection and robust tracking of small moving targets, characterized in that: include: A small target detection network is constructed based on the YOLOv5 network model, wherein the main structure in the small target detection network is the YOLOv5 network structure, wherein the backbone network structure and the prediction head structure of the YOLOv5 network structure are adjusted and a CBAM module is added, and the adjusted YOLOv5 network structure is trained to obtain a target detection network; Acquire image and video data, detect the image and video data through a small target detection network to generate different target detection results, and track the target detection results through a tracking algorithm to generate a target tracking result; Adjusting the backbone network structure includes: The backbone network structure of the YOLOv5 network structure is replaced with the MobileNetV3 network structure, and the original convolution in the backbone network structure is replaced with a lightweight convolution, wherein the lightweight convolution is a GhostConv lightweight convolution; The MobileNetV3 network structure includes an hourglass module, wherein the hourglass module includes two consecutive inverted residual modules connected in series to form an inverted residual structure, wherein the inverted residual structure is used to improve the nonlinear ability of the model; Adjusting the prediction head structure includes: In the prediction head structure, a small target detection layer is added, wherein the backbone network outputs a 640*640 feature map, and the small target detection layer uses a 160*160 feature map to detect targets with pixels larger than 4*4; A CBAM module is added after each convolutional layer in the YOLOv5 network structure, wherein the CBAM module consists of a channel attention mechanism and a spatial attention mechanism.
2. The method according to claim 1, characterized in that During the training process of the adjusted YOLOv5 network structure, the classification loss function is replaced, where the replaced classification loss function adopts the Focal Loss loss function, and the Focal Loss loss function is: FL(p t )=-a t (1-p t ) γ log(p t ) Among them, FL(p t ) represents the loss value of the Focal Loss loss function, α t Represents the weight coefficient related to the sample category, p t It represents the probability that the model predicts that the sample belongs to the correct category, and γ represents the adjustment parameter.
3. The method according to claim 1, characterized in that: During the training process of the adjusted YOLOv5 network structure, the regression loss function is replaced, where the replaced regression loss function is the CIoU intersection-over-union function, and the CIoU intersection-over-union function is: Among them, L CIoU represents the loss value of the CIoU intersection-over-union function, IoU represents the ratio of the intersection and union of two given shapes, ρ represents the Euclidean distance between the center point of the predicted box and the true box, b and b gt They represent the coordinates of the center points of the predicted box and the real box respectively, c represents the diagonal length of the minimum circumscribed rectangle representing the two images, α represents the balance parameter, and v represents the aspect ratio.
4. The method according to claim 1, characterized in that: The tracking algorithm uses the DeepSORT target tracking algorithm.
Citation Information
Patent Citations
Small target detection method and system for improving yolov5 network
CN114373121A
Small target detection system and method based on improved YOLOv5
CN117523267A
Unmanned aerial vehicle detection method and device, computer equipment and storage medium
CN114529714A
Target detection method and system based on YOLOv5s improvement
CN117333857A