Method for reducing annotation difference of multi-person detection frame in fuzzy boundary scene
By converting the detection box into standard coordinates and merging the same category boxes using adjacency tables and depth-first searches, combined with loss function optimization, the labeling inconsistency problem in edge fuzzy scenarios is solved, and the model detection effect and robustness are improved.
Patent Information
- Application Number
- CN202510402069.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
In object detection tasks, especially in edge fuzzy scenarios, it is difficult for annotation personnel to accurately define the target boundaries, resulting in inconsistent labeling results, increasing the difficulty of model training and reducing detection accuracy.
An algorithm is designed to convert all detection boxes into standard coordinates, determine the connection relationship of the annotation boxes of the same category through the adjacency table and the connection factor, merge the relevant boxes using depth-first search, and calculate the minimum external rectangle as a new mark box, which is applied to loss function optimization model training.
It improves the robustness and accuracy of the object detection model in fuzzy boundary scenarios, and at the same time reduces the work burden of labeling personnel.
Smart Images

Figure CN120259805A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, especially a method for reducing the annotation differences of multiple person detection boxes in object detection, and is particularly applicable to processing datasets with blurred edges, such as scenes with unclear boundaries like fire flames and road cracks. Background Art
[0002] In object detection tasks, the quality of annotated data is crucial for the training effect of the model. However, in practical applications, especially when dealing with objects with blurred edges, annotators often have difficulty accurately defining the boundaries of the objects, resulting in inconsistent annotation results. This inconsistency increases the difficulty of model training and reduces the detection accuracy.
[0003] In the prior art, there are two main processing strategies: one is to estimate the uncertain detection box boundaries from a statistical perspective. This method works well when the overall dataset annotation is consistent, but when the boundaries are blurred and the annotations are inconsistent, the estimation results are prone to errors; the other is to measure the difference between the predicted box and the ground truth box based on optimization methods such as CIOU and GIOU based on the intersection over union. Although it can reduce the impact of inconsistent annotations to a certain extent, it is difficult to fundamentally solve the problem of inconsistent annotations caused by blurred boundaries. Summary of the Invention
[0004] The present invention designs a method that can reduce the impact of differences in multiple annotations of an object detection dataset on the model in scenarios with blurred boundaries. This method is applicable to datasets with blurred edges, such as scenes where the boundaries between annotation individuals are blurred like fire flames and road cracks, and requires annotators to follow the principle of fine annotation, that is, when the edge boundaries of two annotation objects are blurred and it is difficult to clearly determine whether they are a single or two independent objects, they should be annotated according to the standard of two objects. The overall process is as Figure 1 shown in the algorithm flowchart.
[0005] The specific implementation process is as Figure 2As shown in the algorithm pseudocode, we will convert all bounding boxes \(b_i\) into standard upper-left coordinates \(x_{i1}, y_{i1}\) and lower-right coordinates \(x_{i2}, y_{i2}\). Then, we create an adjacency list containing \(n\) sub-lists to store the connection relationships between each bounding box and other bounding boxes. Next, by traversing all combinations of bounding boxes, if two bounding boxes are not of the same category, they are directly regarded as having no connection relationship. If they are of the same category, we calculate the Connection Factor (CF) between the two bounding boxes and compare it with the preset threshold \(\alpha\). If it is less than the threshold, they are regarded as having a connection relationship. The method for the Connection Factor is to first calculate the midpoints \(c_1, c_2\) of the two bounding boxes, where \(c_{1x}, c_{1y}\) represent the \(x\) and \(y\) coordinates of point \(c_1\) respectively. Then, we calculate the distance \(D\) between the two center points and the average value \(d\) of the diagonals of the two bounding boxes. If \(D\) is less than the product of the threshold \(\alpha\) and \(d\), it meets the merging requirement and a connection relationship is established. The larger \(\alpha\) is, the easier it is to establish a connection. The size of \(\alpha\) needs to be adjusted according to the dataset and the task. The specific formula is as follows:
[0006] Add its index to the sub-list at the corresponding position in the adjacency list. Finally, use the depth-first search algorithm (DFS) to traverse the connection relationships of each bounding box, store the indices of the bounding boxes belonging to the same component in a list, and add this list to the final component list components. Calculate the minimum bounding rectangle of all rectangles in the same component as its new labeled bounding box, and remove the merged rectangles.
[0007] Then we apply this method to the loss function. Different from the usual loss function, the usual loss function mainly optimizes the prediction results, while we optimize the ground truth. Our method is applied to the bounding box regression function, where our method is called AMA, \(b_{label}\) is the labeled bounding box, \(b_{pred}\) is the predicted bounding box, \(LossFunction\) is the original bounding box regression loss function, and \(Lossreg\) is the optimized bounding box regression loss function. The formula is as follows:
[0008] By applying this algorithm to the regression function, we can use the fused and optimized detection bounding boxes to guide the training of the model, so that the detection bounding boxes of the model have more explicit regression targets, making the regression process of the bounding boxes smoother and ultimately improving the overall detection effect. Description of the Drawings
[0009] Figure 1 This is the flowchart of the method of the present invention, showing the whole process from the overall algorithm implementation to the loss function optimization.
[0010] Figure 2 This is the pseudocode of the method of the present invention, showing the specific implementation of the algorithm. Specific implementation mode
[0011] Input data preparation: Obtain a set of bounding box collections B, where each bounding box bi is represented as (x1i, y1i, x2i, y2i), representing the upper left and lower right coordinates of the bounding box respectively. Set a threshold α to determine whether there is a connection relationship between two bounding boxes.
[0012] Bounding box normalization: Convert all bounding boxes into the standard coordinate form, that is, the x and y coordinates of the upper left point and the lower right point.
[0013] Construct an adjacency list: Initialize an adjacency list A containing n lists to store the connection relationships between each bounding box and other bounding boxes.
[0014] Calculate the connection factor and construct the adjacency relationship: For each pair of bounding boxes (bi, bj), if bi and bj do not belong to the same category, skip this pair of bounding boxes and continue with the next pair. Calculate the connection factor CF(bi, bj). The calculation method of the connection factor includes calculating the midpoints c1, c2 of the two bounding boxes, the distance D between the two center points, and the average value d of the diagonals of the two bounding boxes. If D is less than the product of the threshold α and d, it is considered that there is a connection relationship between the two bounding boxes. If CF(bi, bj) < α, add j to A[i] and add i to A[j], indicating that there is a connection relationship between bi and bj.
[0015] Depth - first search (DFS): Initialize an empty component list C. For each bounding box bi, if it has not been visited, perform the following operations: Initialize an empty component list components, start depth - first search from bi, add all the bounding boxes associated with bi to components, and add components to C.
[0016] Calculate the minimum - enclosing rectangle and update the bounding box set: For each component c in C, calculate the minimum - enclosing rectangle of all the bounding boxes in c, remove the merged rectangles from the original bounding box set B, and add the new bounding boxes to B.
[0017] Apply the optimized bounding box to the loss function: Apply the above algorithm to the bounding box regression function, and guide the training of the model through the optimized bounding box regression loss function Lossreg. The calculation formula of Lossreg takes into account the labeled bounding box blabel, the predicted bounding box bpred, and the original bounding box regression loss function LossFunction.
[0018] Through the above steps, the method of the present invention can intelligently perform automated fusion processing on multiple labeled boxes of the same category, reduce the impact of inconsistent labeling on model training, and improve the detection performance of the model.
Claims
1. A method for reducing the annotation differences of multi-person detection boxes in a fuzzy boundary scenario, characterized in that, Including the following steps: Convert all detection boxes into standard coordinates, and each detection box is represented by the upper-left corner coordinates and the lower-right corner coordinates; Create an adjacency list containing multiple sub-lists to store the connection relationships between each detection box and other detection boxes; Traverse all combinations of detection boxes, calculate the connection factor between detection boxes of the same category, and determine whether two detection boxes have a connection relationship according to a preset threshold; Use the depth-first search algorithm (DFS) to traverse the connection relationships of each detection box, merge the detection boxes belonging to the same component, and calculate the minimum bounding rectangle of the merged detection boxes as the new marked box; Apply the above algorithm to the loss function to guide the training of the model by optimizing the true bounding boxes.
2. The method according to claim 1, wherein The calculation method of the connection factor is to first calculate the midpoints of two detection boxes, then calculate the distance between the two center points and the average value of the diagonals of the two detection boxes, and determine whether the two detection boxes meet the merging requirements according to the product of the preset threshold and the average value.
3. The method according to claim 1 and 2, characterized in that, The loss function is a bounding box regression loss function, and the optimized detection boxes are used to guide the training of the model by fusion, so that the detection boxes of the model have more explicit regression targets.
4. The method according to claim 1, wherein This method is applicable to datasets with blurred edges, such as scenarios where the boundaries between labeled individuals are blurred, such as fire flames and road surface cracks, and requires labelers to follow the principle of fine labeling.
5. The method according to claim 1, characterized in that This method further includes the steps of removing the merged detection boxes and adding the new marked boxes to the detection box set.