A long-tail object detection method based on architecture-agnostic loss
Patent Information
- Application Number
- CN202410803402.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-06-20
AI Technical Summary
然而,现有的方法并未关注到目标检测中目标位置及目标长宽比的长尾分布问题,结果导致目标位置、长宽比或边界框旋转角处于分布尾部的目标检测性能差
(1)本发明提出了一种针对目标位置及边界框具有长尾分布的目标检测任务的一种新的与架构无关的平衡损失,它可以作为插件应用于任何对象检测模型之上,易于实现,能够实现在目标的位置、大小、宽高比、角度变化的情况下对目标的稳定检测;有效利用了目标值的连续性,使得经过平滑后目标值概率分布与预测误差成反比,从而将长尾目标定位的回归问题离散为可进行类别平衡的长尾识别问题。
Smart Images

Figure CN118521774B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection, and more specifically to a long-tail target detection method based on architecture-independent loss. Background Technology
[0002] Long-tailed distribution refers to a distribution pattern where data samples are arranged from largest to smallest frequency according to size or label. In practical applications, training samples often exhibit a long-tailed distribution, where a few classes have a large number of samples, while most other classes have only a small number of samples.
[0003] Long-tail learning can be considered a challenging subtask of class imbalance learning. The main difference lies in the distribution: in long-tail learning, the classes follow a long-tail distribution, which is unnecessary for class imbalance learning. In long-tail learning, the number of classes is typically large, and the tail class samples are often very sparse, whereas in class imbalance learning, the number of minority class samples is not necessarily small in absolute terms. These additional challenges make long-tail learning more challenging than class imbalance learning. Despite these differences, both attempt to address the class imbalance problem.
[0004] Class imbalance increases the training difficulty of deep networks based on recognition models. Trained models tend to favor the head classes, which have abundant training data, while performing poorly on the tail classes, which have limited data. Furthermore, the lack of quantity and diversity of tail class samples makes training models for tail class classification even more challenging. Therefore, deep learning models trained using empirical risk minimization methods generally cannot effectively handle real-world applications with long-tailed class distributions, such as face recognition, species classification, medical image diagnosis, urban scene understanding, and drone detection.
[0005] Many methods have been proposed to address the imbalance problem, including resampling, reweighting, transfer learning, data augmentation, and decoupled training. Among these, reweighting addresses the long-tail distribution problem by improving the loss function, assigning higher weights to the tail classes. This method is simple to implement, and highly competitive results can be achieved by modifying the loss function, making it easily applicable to complex tasks.
[0006] Existing long-tail recognition and long-tail object detection tasks typically focus on class imbalance, favoring datasets with skewed class distributions to improve detection and classification performance for tail classes with small sample sizes. Furthermore, small object detection tasks also address the long-tail distribution of size to enhance detection capabilities for small objects. However, existing methods neglect the long-tail distribution of object location and aspect ratio in object detection, resulting in poor detection performance for objects whose location, aspect ratio, or bounding box rotation angle lies at the tail end of the distribution. Unlike class labels, object center point location, bounding box aspect ratio, size, and bounding box rotation angle are continuous ordinal variables. Therefore, applying existing solutions to class imbalance to address the imbalanced distribution of object center point location, bounding box aspect ratio, size, and bounding box rotation angle in the regression branch is crucial. Summary of the Invention
[0007] The technical problem this invention aims to solve is to provide a long-tail object detection method based on architecture-independent loss, which supplements existing long-tail object detection methods that only focus on class imbalance. It considers the continuity, orderliness, and long-tail distribution characteristics of variables such as the target's center position (x, y), the bounding box's aspect ratio r, size s, and bounding box rotation angle ro. It smooths (x, y, r, s, ro) to align it with the orderliness of the features, and employs the idea of loss reweighting, giving greater weight to bounding boxes located at the tail of the distribution and less weight to bounding boxes located at the head of the distribution, thereby addressing the long-tail distribution problem of (x, y, r, s, ro) in the regression branch.
[0008] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution: A long-tail target detection method based on architecture-independent loss, comprising the following steps: Step (1) uses the labeled publicly available dataset as a basis and performs data augmentation preprocessing using methods including random cropping, scaling, and rotation to obtain the target to be processed; Step (2) Smooth the distribution of the target using Gaussian; Step (3) Construct a feature extraction network to extract features from the input image and obtain the common feature layer of the classification branch and the regression branch; Step (4) Construct the feature layer needed for the regression branch based on the features extracted in step (3); Step (5) Based on the feature-target consistency regularization loss function, the features are aligned with the smoothed target value. Here, the target value is the five-dimensional vector (x, y, r, s, ro) composed of the labeled center point coordinates, the aspect ratio of the bounding box, the size, and the rotation angle. (6) Based on the long-tail distribution characteristics of the target's center point coordinates, the aspect ratio, size, and rotation angle of the bounding box, construct the weights of the target regression loss, give greater weight to the bounding box located at the tail of the distribution, and give less weight to the bounding box located at the head of the distribution. (7) Construct a detection head network to predict the location and category of the target.
[0009] Furthermore, the Gaussian smoothing used in step (2) employs five 1-dimensional Gaussian kernels to smooth each component in (x, y, r, s, ro), where σ is the standard deviation of the Gaussian kernel, which is equivalent to a 5-dimensional Gaussian kernel, with the same standard deviation σ in each direction. For a discrete Gaussian kernel with a window size of w, Let i be the distribution of component i, where Gaussian smoothing is used to smooth the surface. and Perform convolution: , Here, σ specifies the importance assigned to each bin when considering its neighborhood. By increasing σ, the weight assigned to each adjacent bin is increased; the negative correlation between the smoothed target value distribution and the test error distribution is greatly improved, thus enabling the transfer of the distribution-based inverse probability loss weighting method from long-tail classification to the regression branch of long-tail detection.
[0010] Furthermore, in step (3), feature extraction is performed on the features extracted in step (2) to obtain the features required for regression branch prediction.
[0011] Furthermore, the feature extraction backbone network used in step (4) can be of any architecture, because the present invention proposes a new architecture-independent balance loss that can be applied as a plug-in to any object detection model.
[0012] Furthermore, the feature-target consistent ordering regularization loss used in step (5) ensures feature approximation between adjacent targets, and lower feature similarity between target values that are farther apart. That is, in addition to the relationship between adjacent samples involved in Gaussian smoothing, it also considers the ordered relationship between samples with far-away target values, thereby achieving ordered alignment between features and smoothed target values. Feature-target consistent ordering regularization loss The expression is as follows: , Where [i: ] represents the i-th row of matrix S, Indicates the similarity between targets. Indicates the similarity between features. Let represent the rank of the i elements in a; The element at (i, j) is: , A function to describe target similarity, such as cosine similarity; of The element at that location is: , A function that describes feature similarity, such as cosine similarity; This describes the similarity between the i-th objective value and all M objective values. The similarity between the feature corresponding to the i-th target value and the features of all M target values. Mean squared error loss can be used.
[0013] Furthermore, the structure of the regression loss weight construction module used in step (6) is as follows: the target value after Gaussian Amplification is clustered into K classes by the Density-Based algorithm, and the corresponding effective sample number of the K classes is proposed; Consider that each sample occupies a certain volume in the feature space, where S represents the set of all possible data in the feature space for a certain category. Each data point is defined as a subset of S with a volume of 1 unit. Assume the volume of S is N (N≥1). Each data point may overlap with other data points. The more sampled data, the better the coverage of S. The expected volume of the sample increases with the amount of data, eventually reaching the final boundary N. Define the effective data volume of a sample as its expected volume. Considering that there are two relationships between new samples and previously sampled samples during the sampling process: new samples are completely covered by information from old samples, and new samples are not completely covered by information from old samples, the number of effective samples for the K categories proposed in this invention is given by the following formula; , in, This represents the number of samples in the y-th class. This represents the number of valid samples in class y; These are hyperparameters, used to control... The growth rate increases with the number of samples in each class; however, as the number of samples increases, the additional benefits from new data points decrease, potentially leading to a marginal effect. When using heavy data augmentation, newly added samples are likely to be approximate copies of existing samples. Using the effective sample count provides a better characterization of the data. Inverting the effective sample count for each class yields the "class" weights for each regression head: , Therefore, the overall network loss function is obtained: , in, , For classification head categories The number of valid samples, For classification head loss, To recover the head loss, To balance the weights of various losses.
[0014] Compared with existing technologies, the advantages of this invention are: this long-tailed target detection method based on architecture-independent loss considers the long-tailed distribution of target position, size, and angle in target detection, and balances the prediction of this long-tailed distribution, thereby improving the detection of targets with special positions, angles, or small sizes. It has the following advantages: (1) This invention proposes a new architecture-independent balance loss for target detection tasks with long-tailed distribution of target location and bounding box. It can be applied as a plug-in to any object detection model, is easy to implement, and can achieve stable detection of targets under changes in target position, size, aspect ratio, and angle. It effectively utilizes the continuity of target values, so that the probability distribution of target values after smoothing is inversely proportional to the prediction error, thereby discretizing the regression problem of long-tailed target localization into a long-tailed recognition problem that can be balanced by class.
[0015] (2) This invention designs a high-dimensional target value and feature order alignment. By applying the feature target consistent order regularization method applied in age estimation and text similarity research to target detection, the target value: center point coordinates, aspect ratio, size, rotation angle, is aligned with the target feature. That is, the closer the target value is to the sample, the higher the corresponding feature similarity, thus improving the feature extraction capability.
[0016] (3) This invention proposes a new method for calculating the number of valid samples in each class after clustering the target value. The number of valid samples is used to weight the regression loss and balance the performance of the head class and the tail class. Attached Figure Description
[0017] Figure 1 This is an overall flowchart of a long-tail target detection method based on architecture-independent loss according to the present invention; Figure 2 This is a detailed structural diagram of the feature-target consistency ordering regularization method; Figure 3 The results are from testing the VisDrone source domain dataset after training the detector of this invention on that dataset. Figure 4 The results are obtained by testing the detector of this invention on the DOTA source domain dataset after training it. Detailed Implementation
[0018] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, a long-tail target detection method based on architecture-independent loss according to the present invention includes the following steps: Step (1) uses the labeled publicly available dataset as a basis and performs data augmentation preprocessing using methods including random cropping, scaling, and rotation; Step (2) Smooth the distribution of the target value using Gaussian; Step (3) Construct a feature extraction network to extract features from the input image and obtain the common feature layer of the classification branch and the regression branch; Step (4) Construct the feature layer needed for the regression branch based on the features extracted in step (3); Step (5) as follows Figure 2 As shown, based on the feature-target consistency regularization loss function, the features are aligned with the smoothed target value in terms of orderliness. Here, the target value is a five-dimensional vector (x, y, r, s, ro) consisting of the labeled center point coordinates, the aspect ratio of the bounding box, the size, and the bounding box rotation angle. Step (6) Construct the weights of the target regression loss based on the long-tail distribution characteristics of the target's center point coordinates, the aspect ratio, size, and rotation angle of the bounding box. Give the bounding box located at the tail of the distribution a larger weight for the regression loss, and give the bounding box located at the head of the distribution a smaller weight for the regression loss. Step (7) Construct the detection head network to predict the location and category of the target to be detected; Furthermore, the data augmentation methods used in step (1) include random cropping, scaling, and rotation, which are used to alleviate the long-tail characteristics of target position, size, and rotation angle. Random cropping involves randomly selecting a sub-region from the original image, with the region size being 50% to 80% of the total image size, without rescaling the cropped sub-region back to the original image size. This allows the model to learn to adapt to objects of different sizes and positions, improving the model's robustness to size and position changes. Scaling involves randomly adjusting the size of the original image to 0.65 to 1.35 times, improving the model's scale invariance. Rotation involves randomly selecting an angle from the range of [-25°, 25°] to rotate the image, reducing the model's sensitivity to image rotation angles. The Gaussian smoothing used in step (2) employs five 1D Gaussian kernels to smooth each component in (x, y, r, s, ro), where σ is the standard deviation of the Gaussian kernel, equivalent to a 5D Gaussian kernel, with the same standard deviation σ in each direction. Let , For a discrete Gaussian kernel with a window size of w, For the distribution of component i, and Perform convolution to achieve Smoothness: , Here, σ specifies the importance we assign to each bin when considering its neighborhood. By increasing σ, we increase the weight given to each adjacent bin.
[0020] The negative correlation between the smoothed target value distribution and the test error distribution is greatly improved, which facilitates the transfer of the loss weighting method based on the inverse probability of the distribution in long-tail classification to the regression branch of long-tail detection.
[0021] The feature extraction backbone network used in step (3) can be any architecture, such as the ResNet structure, because this invention proposes a new architecture-independent balance loss that can be applied as a plug-in to any object detection model.
[0022] In step (4), based on the features extracted in step (3), further feature extraction is performed through some additional convolutional and fully connected layer structures to predict the category of each target.
[0023] In step (5), such as Figure 2 As shown, the target value The correlation matrix between them is .Target Features are obtained through a feature extraction network. The correlation matrix between features is calculated. The similarity between the target values of the i-th target and all other targets, and the similarity between their corresponding features, are respectively... and Sort the similarities from largest to smallest to obtain... and The rank, that is and Generally, a feature-target consistency ordering regularization loss is constructed. as follows, , in, Let represent the rank of the i elements in vector a. Indicates the similarity between target values. The element at (i, j) is: , A function to describe target similarity, such as cosine similarity; Indicates the similarity between features. The element at (i, j) is: , Let [i: ] be a function describing target similarity, such as cosine similarity, where [i: ] represents the i-th row of matrix S. This describes the similarity between the i-th objective value and all M objective values. The similarity between the feature corresponding to the i-th target value and the features of all M target values. Mean squared error loss can be used.
[0024] Use feature target consistency ordering regularization loss This ensures that samples adjacent to the target value also have a certain degree of similarity in features. The farther apart the target values are, the lower the feature similarity, thus achieving orderly alignment between features and the smoothed target value. The structure of the regression loss weight construction module used in step (6) is as follows: the target values after Gaussian smoothing are clustered into K classes using the Density-Based algorithm, and the effective sample count of the K classes is proposed; Consider that each sample occupies a certain volume in the feature space, where S represents the set of all possible data in the feature space for a certain category. Each data point is defined as a subset of S with a volume of 1 unit. Assume the volume of S is N (N≥1). Each data point may overlap with other data points. The more sampled data, the better the coverage of S. The expected volume of the sample increases with the amount of data, eventually reaching the final boundary N. Define the effective data volume of a sample as its expected volume. Considering that during the sampling process, there are two relationships between new samples and previously sampled samples: the new sample is completely covered by the information of the old sample, and the new sample is not completely covered by the information of the old sample, the number of valid samples in K categories is given by the following formula; , in, This represents the number of samples in the y-th class. This represents the number of valid samples in class y; Hyperparameters, controlling The rate of increase with the number of samples in each category; however, as the number of samples increases, the additional benefits from new data points decrease, potentially leading to a marginal effect. When using data augmentation, newly added samples are likely to be approximate copies of existing samples; using the effective number of samples provides a better characterization of the data. Inverting the effective number of samples yields the weight of each "category" for each regression head; The overall loss function of the network is as follows; , in, For the classification head loss, Focal Loss can be used. and The same definition applies; the weight is the inverse of the number of valid samples in category c in the classification header. , For regression head loss, IOU Loss can be used. For example... Figure 1 As shown, the bounding box loss for target i is The bounding box loss for all n targets in a given image is... Weighted, The regression loss for a given image is obtained by assigning the regression loss weights to the "category" to which i belongs: , By using the number of valid samples for each category in the classification head and regression head, a loss function independent of the network architecture is constructed. This method is simple and easy to implement.
[0025] Figure 3 and Figure 4 The results of training the detector of this invention on the VisDrone and DOTA source domain datasets are shown respectively. The brighter boxes represent newly detected targets after adding regression weights and regularization loss to the ReDet network. It can be seen that the method of this invention improves the detection performance of tail categories, such as small targets with rotation angles at the tail of the distribution.
[0026] It should be emphasized that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A long-tail target detection method based on architecture-independent loss, characterized in that, Includes the following steps: Step (1) Using the labeled publicly available dataset as a base, perform data augmentation preprocessing using methods including random cropping, scaling, and rotation to obtain the labeled targets to be processed; Step (2) Gaussian smoothing is applied to the distribution of the center point position, aspect ratio, size, and rotation angle of the bounding box of the labeled target. Step (3) Construct a feature extraction network to extract features from the input image, obtain the common feature layer of the classification branch and the regression branch, and use the feature layer to extract features; Step (4) Based on the features extracted in step (3), construct the feature layer needed for the regression branch and further extract features; Step (5) Based on the regularized loss function that maintains the consistency and order of features and targets, the features extracted in step (4) are aligned with the order of the smoothed target values. Here, the target value is a 5-dimensional vector (x, y, r, s, ro) consisting of the center point coordinates of the labeled target, the aspect ratio of the bounding box, the size, and the rotation angle of the bounding box. Step (6) Construct the weights of the target regression loss based on the long-tail distribution characteristics of the target's center point coordinates, the aspect ratio of the bounding box, its size, and the bounding box rotation angle. Give the bounding box located at the tail of the distribution a larger weight for the regression loss, and give the bounding box located at the head of the distribution a smaller weight for the regression loss. Step (7) Construct the detection head network to predict the location and category of the target to be detected; The feature-target consistent ordering regularization loss used in step (5) achieves the orderly alignment of features with smoothed target values by making the features similar between adjacent target values, while the feature similarity between target values that are farther apart is lower; feature-target consistent ordering regularization loss The expression is as follows: , Where [i: ] represents the i-th row of matrix S, ∈ Indicates the similarity between targets. ∈ Indicates the similarity between features. Let represent the rank of the i elements in a; The element at (i, j) is: , A function describing target similarity; The element at (i, j) is: , This is a function that describes feature similarity. This describes the similarity between the i-th objective value and all M objective values. The similarity between the feature corresponding to the i-th target value and the features of all M target values. Mean squared error loss is used.
2. The long-tail target detection method based on architecture-independent loss according to claim 1, characterized in that, The Gaussian smoothing used in step (2) employs five 1-dimensional Gaussian kernels to smooth each component in (x, y, r, s, ro), where i ∈ {x, y, r, s, ro}. This represents a discrete Gaussian kernel with a window size of w. Let be the discrete distribution form of component i, that is, the distribution of i is discretized into m bins, σ is the standard deviation of the Gaussian kernel, and the same standard deviation σ is used in all five directions. σ specifies the importance given to each bin when considering the neighborhood. By increasing σ, the weight given to each adjacent bin is increased. and Perform convolution operations: , The negative correlation between the smoothed target value distribution and the test error distribution is improved, thus allowing the regression loss of long-tail target detection to be weighted based on the inverse probability of the distribution.
3. The long-tail target detection method based on architecture-independent loss according to claim 2, characterized in that, In step (3), further feature extraction is performed based on the features extracted in step (2) to obtain the features required for regression branch prediction.
4. The long-tail target detection method based on architecture-independent loss according to claim 1, characterized in that, The feature extraction backbone network used in step (4) can be of any architecture.
5. The long-tail target detection method based on architecture-independent loss according to claim 1, characterized in that, The structure of the regression loss weight building module used in step (6) is as follows: the target value after Gaussian smoothing is clustered into K classes by the Density-Based algorithm, and the effective sample number of the K classes is proposed; Consider that each sample occupies a certain volume in the feature space. S represents the set of all possible data in the feature space of a certain category. Each data is defined as a subset of S with a volume of 1 unit. Suppose that the volume of S is N, where N≥1. Each data may overlap with other data. The more data sampled, the better the coverage of S. The expected volume of the sample will increase with the increase of the amount of data and reach the final boundary N. Define the effective amount of data of the sample as the expected volume of the sample. Considering the two relationships between new and previously sampled samples during the sampling process: the new sample is completely covered by the information of the old sample, and the new sample is not completely covered by the information of the old sample, the effective number of samples for each of the K categories is proposed as follows: , in, This represents the number of samples in the y-th class. This represents the number of valid samples in class y; α∈[1,2) is a hyperparameter that controls... The growth rate increases with the number of samples in each category; however, as the number of samples increases, the additional benefits from new data points decrease, exhibiting a marginal effect. When using heavy data augmentation, newly added samples are likely to be approximate copies of existing samples. Using the effective number of samples provides a better characterization of the data. Inverting the effective sample count yields the weight of the "class" for each regression head; , Overall network loss function: , in, The number of valid samples of category c in the classification head. The weights formed by taking the inverse, , in, To balance the weights of the various losses, For classification head loss, This is a loss due to regression.
Citation Information
Patent Citations
Target detection method based on long-tail distribution data set
CN111898685A
Long-tail target detection method and system
CN113989519A