Efficient rotating target detection method

By adding the scale detection part and improving the loss function in YOLOv5's rotation object detection algorithm, the problem of low accuracy in rotation object detection and difficult for model to converge quickly is solved, and a more efficient rotation object detection accuracy is achieved.

CN119942056APending Publication Date: 2025-05-06YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311452390.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems with low accuracy and difficult model to converge quickly in rotational object detection, especially when detecting objects at rotation angles in remote sensing images.

Method used

Based on YOLOv5's rotation object detection algorithm, by adding a scale detection part after YOLOv5's backbone, the features of the small target are extracted, and the coordinate offset rotation object prediction is performed on the feature map. Combining YOLOv5's loss function and offset loss function, the rotation object is processed after the rotation object. At the same time, the horizontal box loss function of YOLOv5 is improved as the loss function of the rotation box, and a mixed loss function is used, including the IoU loss of the vertical box corresponding to the rotation box and the norm loss of the rotation angle.

Benefits of technology

The accuracy of rotational small object detection is improved, and the problems of angle mutation and edge exchange are solved. By limiting the angle range of the rotation box and using a mixed loss function, the ambiguity problem is avoided, and the accuracy and stability of the detection model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942056A_ABST
    Figure CN119942056A_ABST
Patent Text Reader

Abstract

The invention discloses a rotating target detection method based on high efficiency. The method has certain universality in the rotating target detection direction, and the patent takes remote sensing image target detection as an explanatory case. Based on a baseline network, the network is improved. Firstly, starting from a backbone network for feature extraction, the backbone network with better feature extraction capability and higher efficiency and a multi-scale structure are used, so that the network is enhanced in the aspect of feature extraction. And then various data enhancement methods are used to improve the detection generalization ability of the network. Finally, the invention provides a mixed loss applied to rotating target detection and a rotating frame representation method matched with the mixed loss, so that the network can have better adaptability to regression of the rotating frame. In combination with the method, an efficient rotating target detection network is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a rotation target detection method based on a deep neural detection network YOLOv5, and the invention belongs to the field of computer application. Background Art

[0002] Images are an important source for humans to understand the world, and they can convey richer, more real and specific information than other forms. With the continuous development of social economy and the acceleration of urbanization, the order of cities seems to be more and more disordered. As an important branch of computer vision (CV), object detection is being widely used in industrial inspection, road traffic, aerospace and other fields. For example, using cameras to capture, process and store information about people on the road in real time, so as to reduce the intensity of criminal investigation and reduce the consumption of human capital. Therefore, this technology has important practical significance. According to the direction of the target frame, object detection can usually be divided into horizontal detection and rotation detection. Specifically, horizontal frame detection is usually more suitable for general natural scene images. However, scenes such as remote sensing images, face recognition and license plate recognition usually require more accurate positioning, which requires an effective rotation target detection model. The YOLO (You Only Look Once) series of algorithms are well-known one-stage detectors for detecting horizontal targets. The detector regresses the positioning and classification of targets together. After the image passes through the YOLO backbone network, the position and category of each target object will be directly output. Finally, the corresponding algorithm is used to remove overlapping boxes and perform other post-processing operations. A more representative network is YOLOv3, which has made some improvements on its predecessors YOLOv1 and YOLOv2, such as adding an anchor mechanism, mainly referring to the design of the Feature Pyramid Network (FPN), and can perform multi-scale detection on the input image. SCRDet can be used to solve the problem of rotation detection. The model regresses five parameters, namely the coordinates, width, height and rotation angle of the center point, to describe the rotation bounding box. In order to more accurately predict the rotation box, SCRDet adds an IoU constant factor to the smooth L1 loss function. Due to the inherent periodicity of the angle, the loss discontinuity caused by the sudden exchange of the width and height of the target, and the difference in coordinate and angle units, simply considering the coordinates and angles in the five-parameter system together will lead to training instability and performance degradation. Another rotation target detector RSDet uses an eight-parameter regression method that can use the same unit coordinates to alleviate this problem, further solving the problem of inconsistent parameter regression. In the P-RSDet model, target detection at any angle can be achieved by predicting the center point and regressing a polar radius and two polar angles. In addition, in order to express the geometric constraint relationship between the polar radius and the polar angle, the model uses a polar ring area loss function to improve the prediction accuracy. P-RSDet achieves better performance with a simpler model and fewer regression parameters. This patent focuses on solving the problem of detecting targets that are rotated at a certain angle.In some specific scenarios, such as satellite remote sensing image detection, the targets in the image generally have rotation angles. At this time, if conventional target detection methods are used to detect objects in the image, the accuracy will be low. If a smaller rectangular box can be used to mark these objects, the detection model will be more focused on the target object, thereby solving the low accuracy problem caused by the background around the target. One-stage target detection models like YOLO have some defects, such as the inability to detect small objects well. This is because the target category and position are regressed at the same time, and the model cannot converge quickly. Therefore, for the problem of target detection of dense small objects, how to improve the detection accuracy of the model has certain research value. Summary of the invention

[0003] 1. A rotating target detection algorithm based on YOLOv5, comprising the following steps:

[0004] (1) receiving an input image, and the data loading module performs data enhancement on the input image;

[0005] (2) Add a scale detection part after the backbone of YOLOv5, such as Figure 1 As shown, features are extracted for small targets;

[0006] (3) Predicting the coordinate offset rotation target on the extracted feature map;

[0007] (4) Combine the loss function of YOLOv5 during training and add the offset loss function;

[0008] (5) The predicted target frame is subjected to post-processing of rotating the target, and finally the detection result is obtained;

[0009] (6) Make a rotation anchor target prediction on the extracted feature map;

[0010] (7) Improve the horizontal box loss function of YOLOv5 during training and change it to the loss function of the rotation box;

[0011] (8) The predicted target frame is post-processed by rotating the target and finally the detection result is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 Figure 1: Schematic diagram of the backbone network structure.

[0013] Figure 2 Schematic diagram of the rotating box.

[0014] Figure 3 : Schematic diagram of rotating box mixing. DETAILED DESCRIPTION

[0015] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0016] This paper proposes a rotating target detection network based on improved yolov5. This network mainly improves the detection accuracy of rotating small targets in current remote sensing scenes. The network structure is shown in the figure Figure 1 The purpose of the present invention is to improve the following aspects:

[0017] 1. For the detection of rotating targets using the traditional five-parameter representation, using the norm loss as the rotation box loss will cause the problems of angle mutation PoA and edge exchangeability EoE. The root of this problem lies in the periodicity of the angle. Due to the existence of the angle periodicity, a rotation box may have multiple representations that are different in value but the same in practical sense, that is, ambiguity. For example, [100, 100, 30, 40, 40°] and [100, 100, 40, 30, -50°] actually point to the same rotation box. Therefore, in order to solve the ambiguity problem, the present invention uses a method similar to the OpenCV representation to limit the angle range of the rotation box to (-45°, 45]. In the network output part, the angle output of the network is limited to a certain range to avoid ambiguity. The network will not produce output values ​​with huge differences, while the actual rotation boxes are similar. The rotation box representation used in the present invention is as follows: Figure 2 Shown

[0018] 2. Since there are too many difficulties in directly using the rotation IoU loss, in order to take advantage of the excellent characteristics of the IoU loss, this paper proposes a hybrid loss, which splits the regression loss of the rotation box into the IoU loss of the vertical box corresponding to the rotation box and the norm loss of the rotation angle. That is, a rotation box is mixed and the loss is calculated. This operation is as follows Figure 3 shown.

[0019] Step 1: Modify the backbone network to CSPNet of YOLO-v5 and use PAFPN for multi-scale feature fusion.

[0020] Step 2: Add strong data enhancement including mosaic enhancement in the training stage, use the hybrid regression loss proposed in this invention as the regression loss for training, and obtain an efficient rotation object detection network

[0021] Step 3: The loss function of the network is: L = λL reg +L cls +L mask

[0022] Step 4: The regression loss is L reg =αL IoU +L angle The vertical box IoU loss L IoUUsing the form of 1-IoU, the angle regression loss L angle Use Smooth-l1 loss. α is the balance coefficient to balance the two losses.

[0023] Step 5: The CSPNet used in the network has stronger feature extraction capabilities than ResNet, and the overall computational complexity is effectively controlled, and the network speed is improved; at the same time, the addition of PAFPN enables the network to have better detection performance for remote sensing targets with a larger scale distribution range. With the use of multiple data enhancement methods and the hybrid loss proposed in this invention, the detection capability of the network is greatly enhanced.

Claims

1. A rotating target detection algorithm based on YOLOv5, characterized in that: The following steps are involved: Step 1: receiving an input image, and the data loading module performs data enhancement on the input image; Step 2: Add a scale detection part after the backbone of YOLOv5, as shown in Figure 1, to extract features for small targets; Step 3: Predict the coordinate offset rotation target on the extracted feature map; Step 4: Combine the loss function of YOLOv5 with the offset loss function during training; Step 5: Rotate the predicted target frame and then get the detection result. Step 6: Make a rotation anchor target prediction on the extracted feature map; Step 7: Improve the horizontal box loss function of YOLOv5 during training and change it to the loss function of the rotation box; Step 8: Rotate the predicted target frame and then get the detection result.