On-device small target detection method and device based on tracking and adaptive image slicing
By introducing technical means such as adaptive map cutting, heat map and motion trend judgment into the small object detection algorithm, the shortcomings of the existing small object detection algorithm in terms of detection accuracy and complexity are solved, and a more efficient and accurate small object detection effect is achieved.
Patent Information
- Application Number
- CN202311152682.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-09-08
AI Technical Summary
The existing small object detection algorithms have problems with low detection accuracy and high algorithm complexity when processing small objects, especially when the small objects are difficult to distinguish after image scaling, and the sliding window scanning method is time-consuming and has no obvious effect.
The small object detection method on the end based on tracking and adaptive map cutting is adopted. The object area distribution map on the original image is self-learning, and the target size range is judged by combining the heat map and motion trend, and the detection efficiency and accuracy are improved through the dynamic image segmentation and repetitive area suppression module.
It realizes the reduction of algorithm complexity while maintaining detection accuracy, improves the efficiency and accuracy of small object detection, reduces redundant prediction results, and enhances the stability and efficiency of the system.
Smart Images

Figure CN117392169B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and in particular relates to a method and device for detecting small targets on a terminal based on tracking and adaptive graph slicing. Background Art
[0002] How to quickly and accurately locate objects of interest in visible light scenes has been a recent research hotspot, which directly affects the implementation and feasibility of algorithms. In the early days, there were two-stage RCNN series positioning algorithms. The first step was to extract feature maps from the original image through CNN, extract region proposals based on anchor prediction, and the second step was to perform classification and precise position regression based on proposals. Later, there were single-stage SSD series and YOLO series algorithms, which extracted feature maps through CNN deep convolutional networks, directly predicted object categories, and regressed precise position boxes.
[0003] The final detection accuracy of the above two algorithms is directly related to the featuremap representation ability. The feature extraction module has become more complex from the original deep convolutional CNN network structure, adding FPN (feature pyramid), residual building blocks, PAFPN, CSPResNet, and transformer modules, which have increased the algorithm complexity while improving the feature representation ability. In order to improve network performance, deep separable convolution, self-learning modules, and image scaling modules have been added, but while improving performance, the algorithm accuracy has also been lost, especially when the scaling factor is too high. The area of small targets is originally small, and the human eye cannot distinguish them after scaling. Recently, a single-stage algorithm has also been used to scan the original image with a fixed ratio sliding window, and a small image prediction method has been used to compensate for the scaling loss, but the sliding window step size and size are too large, the effect is not obvious, and if it is too small, the time consumption is seriously increased. Summary of the invention
[0004] In view of this, the present invention provides an on-device small target detection method based on tracking and adaptive image slicing, comprising:
[0005] Step 1: The original large image is divided into multiple sub-images according to a fixed ratio and step size, and there is a certain overlap between the sub-images; adaptive image cutting is performed by learning the area distribution map of objects appearing at each position on the original image;
[0006] Step 2, judging the size range of the target based on the heat map of the target's location, the probability of the target appearing at the location, and the target's movement trend;
[0007] Step 3, tracking the dynamic image of the small target, determining the movement trend of the small target, and segmenting the dynamic image of the small target according to the movement trend;
[0008] Step 4: for the multiple prediction boxes where the small target exists, determine whether there is repeated prediction; if at least one of the prediction boxes is a subbox of other prediction boxes, use the area size of the picture where the prediction box where the small target appears as a threshold to suppress the targets in the overlapping areas of the multiple prediction boxes.
[0009] In particular, judging the size range of the target according to the heat map of the target location, the probability of the target appearing at the location, and the movement trend of the target in step 2 specifically includes: judging the size range of the target according to the following formula:
[0010] ;in, is the target size evaluation value calculated according to the given conditions, The heat value at the position (x, y) obtained according to the heat map accumulated in the first step; is the absolute value of the ratio of the probability values of the target at position (x, y); where, represents the target probability value at position (x, y), Represents the target probability value in the entire image; Parameters used for adjustment, indicating the size of the slice or the number of cropped sub-pictures; is the height change of the target at position (x, y); is the probability value of the target at position (x, y), is the average entropy value of the heat map; It is the average value of the evaluation value obtained based on the target area change statistics of the tracking module.
[0011] In particular, judging the movement trend of the small target in step 3 includes: establishing a complete tracking chain for each small target entering and leaving the field of view, and during the tracking process, monitoring the area change of each object is achieved through the following formula:
[0012]
[0013] in, Indicates the maximum area within the time range of the tracking chain, represents the minimum value of the area within the time range in the tracking chain, β represents the weight factor, represents the absolute value of the change in size or area between adjacent time steps; t represents the index of the time step, from 1 to N; in this formula, the motion trend score of the small target is obtained by calculating the sum of the absolute values of the area change of the small target between adjacent time steps and the difference between the maximum and minimum values of the area of the small target within the observation time range. , the rating It can be used to measure the change in size or area of a small target within the observation time range, so as to determine its movement trend; if the movement trend of the small target is from a small area to a large area, the number of cut sub-images is reduced and the sub-image size is increased; if the movement trend of the small target is from a large area to a small area, the number of cut sub-images is increased and the sub-image size is reduced.
[0014] In particular, in step 4, if at least one of the prediction boxes is a subbox of other prediction boxes, based on the area size of the picture where the prediction box where the small target appears is located as a threshold, suppressing the targets in the overlapping areas of multiple prediction boxes specifically includes: sorting multiple prediction boxes with an area greater than the preset threshold according to the area of the preset threshold from large to small according to the area; selecting the box with the largest area as the candidate box; calculating the degree of overlap between the candidate box and the multiple prediction boxes, if the degree of overlap between a certain prediction box and the candidate box is greater than the preset threshold, deleting the boundary box, otherwise retaining it; and continuing to execute until all prediction boxes are processed.
[0015] In particular, the step 4 also includes:
[0016] If there are at least two prediction boxes that are complete boxes including the small target, all prediction boxes are sorted in descending order according to the confidence, and the box with the highest confidence is selected as the candidate box. The degree of overlap between the candidate box and the multiple prediction boxes is calculated. If the degree of overlap between a prediction box and the candidate box is greater than a preset threshold, the bounding box is deleted, otherwise it is retained; continue to execute until all prediction boxes are processed.
[0017] In particular, when calculating the overlap between the candidate box and the multiple prediction boxes by using the intersection-and-union ratio, the intersection-and-union ratio loss function is used. Applied in the model, the intersection-over-intersection loss function is:
[0018]
[0019] in, is the intersection-over-union ratio of the prediction box; is the target value of the intersection-union ratio.
[0020] In particular, by enhancing the intersection-over-union score Calculate the overlap between the candidate box and the multiple prediction boxes, and the enhanced intersection-over-union score Specifically include:
[0021]
[0022] The formula consists of three terms. The first term measures the distance between the center of the predicted box and the center of the candidate box. , relative to the size of the selection box, Represents the width of the selection box. Represents the height of the candidate box; the second term measures the difference in width between the predicted box and the candidate box, relative to the square of the width of the candidate box; Represents the difference in width between the predicted box and the candidate box; the third term measures the difference in height between the predicted box and the candidate box, relative to the square of the height of the candidate box; Represents the height difference between the predicted box and the candidate box.
[0023] In particular, the step 4 also includes: if there is overlap of the small targets in the adjacent multiple prediction frames, merging the upper and lower detection frames or the left and right detection frames according to the overlap ratio.
[0024] The present invention also proposes an on-device small target detection device based on tracking and adaptive image slicing, comprising:
[0025] The self-learning module is used to split the original large image into multiple sub-images according to a fixed ratio and step size, with a certain overlap between the sub-images; adaptive image splitting is performed by learning the area distribution map of objects appearing at each position on the original image;
[0026] A motion trend judgment module is used to judge the size range of the target based on the heat map of the target's location, the probability of the target appearing at the location, and the target's motion trend;
[0027] A dynamic image segmentation module is used to track the dynamic image of the small target, determine the movement trend of the small target, and segment the dynamic image of the small target according to the movement trend;
[0028] The repeated area suppression module is used to determine whether there is repeated prediction for multiple prediction boxes where the small target exists; if at least one of the prediction boxes is a subbox of other prediction boxes, the target in the overlapping area of the multiple prediction boxes is suppressed based on the area size of the picture where the prediction box where the small target appears is located as a threshold.
[0029] Beneficial effects:
[0030] 1. Adaptive image cutting: By learning the area distribution map of objects appearing at each position on the original image, adaptive image cutting of the large image can be achieved. This can effectively divide the position and size of the sub-image according to the density and distribution of objects at different positions, thereby improving the effect of small target detection.
[0031] 2. Motion trend judgment: By analyzing the thermal map of the target's location, the probability of the target's appearance and the motion trend, the size range of the target can be judged. This technical effect can infer the size range of the target based on the location and thermal information of the target in the image, further improving the accuracy of small target detection.
[0032] 3. Dynamic image segmentation: By tracking the dynamic image of small targets and segmenting them according to the movement trend, effective processing of small targets can be achieved. This technical effect can segment the dynamic image into appropriate parts according to the movement trend of the target, so as to better perform subsequent processing and analysis, and improve the efficiency of small target detection.
[0033] 4. Repeated area suppression: By judging the relationship between prediction boxes and the size of overlapping areas, repeated prediction boxes can be suppressed. This technical effect can reduce redundant prediction results and improve the accuracy and precision of small target detection.
[0034] The entire technical solution of the present invention comprehensively applies strategies such as adaptive image cutting, motion trend judgment and repeated area suppression, and improves the comprehensive performance of small target detection through the collaborative work of various modules. This technical effect can make the system more stable, accurate and efficient when processing small targets on the terminal. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flow chart of the on-device small target detection method based on tracking and adaptive image slicing in the present invention;
[0036] Figure 2 Schematic diagram of an on-terminal small target detection device based on tracking and adaptive image slicing in the present invention. DETAILED DESCRIPTION
[0037] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0038] The present invention provides an on-device small target detection method based on tracking and adaptive image slicing, such as Figure 1 As shown, including:
[0039] Step 1: The original large image is divided into multiple sub-images according to a fixed ratio and step size, and there is a certain overlap between the sub-images; adaptive image cutting is performed by learning the area distribution map of objects appearing at each position on the original image;
[0040] In the initial self-learning stage of the algorithm, most algorithms currently scale the original large image to a fixed-size small image in the first step. Some small targets (such as ,even ) after scaling ( even *2) It is difficult to distinguish with the naked eye and difficult for the algorithm to recall. In the initialization stage of the algorithm, the image scaling ratio is reduced to avoid the loss of small targets after scaling. The original image will be divided into several sub-images according to a fixed ratio and step size. At the same time, in order to prevent the breakage of objects in critical areas, each sub-image has a fixed ratio of overlap, and as many small target objects as possible in the original image are recalled, and the area distribution map of objects appearing at each position on the original image is self-learned.
[0041] Step 2, judging the size range of the target based on the heat map of the target's location, the probability of the target appearing at the location, and the target's movement trend;
[0042] In step 2, judging the size range of the target according to the heat map of the target location, the probability of the target appearing at the location, and the movement trend of the target specifically includes: judging the size range of the target according to the following formula:
[0043]
[0044] in, is the target size evaluation value calculated according to the given conditions, The heat value at the position (x, y) obtained according to the heat map accumulated in the first step; is the absolute value of the ratio of the probability values of the target at position (x, y); where, represents the target probability value at position (x, y), Represents the target probability value in the entire image; Parameters used for adjustment, indicating the size of the slice or the number of cropped sub-pictures; is the height change of the target at position (x, y); is the probability value of the target at position (x, y), is the average entropy value of the heat map; It is the average value of the evaluation value obtained according to the target area change statistics of the tracking module. In step 2 of this embodiment, the target probability ratio reflects the relative probability size of the target at the position by calculating the absolute value of the ratio of the target probability value at the position (x, y) to the target probability value in the entire image. A larger ratio indicates that the probability of the target at this position is higher. The height change value is the height change value of the target at the position (x, y), taking into account the size change of the target in the vertical direction. A larger height change value indicates that the vertical size of the target at this position is larger. The magnitude and range of the target size evaluation value are adjusted by adjusting the parameters and the average entropy value of the heat map.
[0045] Step 3, track the dynamic image of the small target, determine the movement trend of the small target, and segment the dynamic image of the small target according to the movement trend; in step 3, first obtain the original image sequence, including continuous frames of multiple time steps, and use the target tracking algorithm, such as Kalman filter, correlation filter or deep learning tracker, to track the small target in the initial frame. The tracker will output the position information of the small target at each time step; according to the output of the tracker, establish the tracking chain of the small target and record the position information of the small target at each time step.
[0046] Determining the movement trend of the small target includes: establishing a complete tracking chain for each small target entering and leaving the field of view, and during the tracking process, monitoring the area change of each object is achieved through the following formula:
[0047]
[0048] in, Indicates the maximum area within the time range of the tracking chain, represents the minimum value of the area within the time range in the tracking chain, β represents the weight factor, represents the absolute value of the change in size or area between adjacent time steps; t represents the index of the time step, from 1 to N; in this formula, the motion trend score of the small target is obtained by calculating the sum of the absolute values of the area change of the small target between adjacent time steps and the difference between the maximum and minimum values of the area of the small target within the observation time range. , the rating It can be used to measure the change in size or area of a small target within the observation time range, so as to determine its movement trend; if the movement trend of the small target is from a small area to a large area, then reduce the number of cut sub-images and increase the sub-image size; if the movement trend of the small target is from a large area to a small area, then increase the number of cut sub-images and reduce the sub-image size. For example, suppose we have a small target tracking chain containing 5 time steps (t1, t2, t3, t4, t5), and we have recorded the area of the small target s(t) at each time step. Calculate the absolute value of the area change between each adjacent time step, where t is 1 to 4. Substituting into the formula, if If the size of the small target is larger, it means that the area of the small target has a significant trend of increasing within the observation time range. It can be inferred that the small target may have entered from a distance and gradually increased. In this case, you can consider reducing the number of cropped sub-images and increasing the sub-image size to better capture the details and features of the small target.
[0049] if If the size of the small target is smaller, it means that the area of the small target changes little or tends to be stable within the observation time range. It can be inferred that the small target may be close and remain relatively stable. In this case, you can consider increasing the number of cropped sub-images and reducing the sub-image size to better capture the shape and motion trajectory of the small target.
[0050] Step 4: for the multiple prediction boxes where the small target exists, determine whether there is repeated prediction; and suppress the targets in the overlapping areas of the multiple prediction boxes.
[0051] At this time, there are three cases. The first case is that in step 4, if at least one of the prediction boxes is a subbox of other prediction boxes, the area size of the picture where the prediction box where the small target appears is located is used as a threshold, and the targets in the overlapping areas of multiple prediction boxes are suppressed, specifically including: according to the area of a preset threshold, sorting multiple prediction boxes with an area larger than the preset threshold from large to small according to the area; selecting the box with the largest area as the candidate box; calculating the degree of overlap between the candidate box and the multiple prediction boxes, if the degree of overlap between a certain prediction box and the candidate box is greater than the preset threshold, deleting the boundary box, otherwise retaining it; continuing to execute until all prediction boxes are processed.
[0052] In the second case, if there are at least two prediction boxes that are complete boxes that include the small target, all prediction boxes are sorted in descending order according to the confidence, and the box with the highest confidence is selected as the candidate box. The degree of overlap between the candidate box and the multiple prediction boxes is calculated. If the degree of overlap between a prediction box and the candidate box is greater than a preset threshold, the bounding box is deleted, otherwise it is retained; continue to execute until all prediction boxes are processed. We have the following three boxes and their confidences:
[0053] Box 1: (x1, y1, x2, y2) = (50, 50, 100, 100), with a confidence level of 0.9
[0054] Box 2: (x1, y1, x2, y2) = (60, 60, 120, 120), with a confidence level of 0.8
[0055] Box 3: (x1, y1, x2, y2) = (70, 70, 90, 90), with a confidence level of 0.7
[0056] Use the non-maximum suppression (NMS) algorithm with confidence as the criterion, for example:
[0057] Sort all boxes in descending order by confidence:
[0058] Box 1: Confidence level 0.9
[0059] Box 2: Confidence level 0.8
[0060] Box 3: Confidence level 0.7
[0061] Select the box with the highest confidence (box 1) and add it to the final list of selected boxes.
[0062] Calculate the intersection over union (IoU) of box 2 and box 1:
[0063] IoU(box 2, box 1) = calculate the intersection area of the two boxes / calculate the union area of the two boxes
[0064] If IoU(box 2, box 1) is greater than the set threshold (for example, 0.5), box 2 is discarded; otherwise, box 2 is added to the final selected box list.
[0065] Calculate the intersection over union (IoU) of box 3 and box 1:
[0066] IoU(box 3, box 1) = calculate the intersection area of the two boxes / calculate the union area of the two boxes
[0067] If IoU(box 3, box 1) is greater than the set threshold (e.g. 0.5), box 3 is discarded; otherwise, box 3 is added to the final selected box list.
[0068] Regarding how to determine the degree of overlap between the candidate box and the multiple prediction boxes, in this embodiment, when calculating the degree of overlap between the candidate box and the multiple prediction boxes by using the intersection-over-union ratio, the intersection-over-union ratio loss function is used. for:
[0069]
[0070] in, is the intersection-over-union ratio of the prediction box; is the target value of the intersection-union ratio.
[0071] The formula is based on the IOU of the predicted bounding box and the target value The comparison is divided into two cases for calculation. If the IOU of the predicted bounding box is less than the target value , then use the first formula to calculate IOU Loss. In this case, is the IOU of the predicted bounding box and the target value If the IOU of the predicted bounding box is greater than or equal to the target value , then use the second formula to calculate .in this case, is the target value The negative of the natural logarithm of the IOU with the predicted bounding box.
[0072] The intersection-over-union loss is used to measure the difference between the predicted bounding box and the target bounding box. When the consistency is good, the intersection-over-union loss approaches 0, indicating that the predicted bounding box has a high degree of overlap with the target bounding box. When the difference between is large, the intersection-over-union loss increases, indicating that the overlap between the predicted bounding box and the target bounding box is low.
[0073] The intersection-over-union loss function can be used to train the target detection model. In the target detection task, the model needs to learn the correspondence between the predicted box and the true labeled box. By using the intersection-over-union loss function, the overlap between the predicted box and the true labeled box can be measured, thereby guiding the model to accurately locate the target and train the bounding box regression.
[0074] In addition, the intersection-over-union loss function can be used as one of the indicators for evaluating the performance of the target detection model. The intersection-over-union ratio between the predicted box and the true labeled box is calculated to measure the accuracy and recall of the model's prediction. Generally, a higher intersection-over-union ratio indicates that the model has better positioning and detection capabilities.
[0075] In this embodiment, when calculating the overlap between the candidate box and the multiple prediction boxes, the intersection-over-union score can be enhanced. Calculate the overlap between the candidate box and the multiple prediction boxes, and the enhanced intersection-over-union score Specifically include:
[0076]
[0077] The formula consists of three terms. The first term measures the distance between the center of the predicted box and the center of the candidate box. , relative to the size of the selection box, Represents the width of the selection box. Represents the height of the candidate box; the second term measures the difference in width between the predicted box and the candidate box, relative to the square of the width of the candidate box; Represents the difference in width between the predicted box and the candidate box; the third term measures the difference in height between the predicted box and the candidate box, relative to the square of the height of the candidate box; Represents the height difference between the predicted box and the candidate box.
[0078] By combining these components, the enhanced intersection-over-union metric can comprehensively consider the position, size, and shape information of the bounding box, providing a more accurate bounding box matching metric. It can be used as a metric for model performance evaluation in object detection tasks, or as a loss function in bounding box regression tasks to help the model learn accurate bounding box predictions.
[0079] There is also a third situation, if the small target overlaps in the adjacent multiple prediction frames, the upper and lower detection frames or the left and right detection frames are merged according to the overlap ratio. In this way, when the large image is cut into sub-images, although there is a certain overlap in the upper and lower and left and right ratios (such as 20% width and height of the sub-image), there are still objects that will be cut off from the upper and lower and left and right. At this time, by calculating the horizontal overlap ratio of the left and right sub-images, the left and right are merged when it is higher than the threshold, and the upper and lower sub-images are also merged in the height direction.
[0080] The present invention also proposes an on-device small target detection device based on tracking and adaptive image slicing, such as Figure 2 As shown, the device comprises:
[0081] The self-learning module is used to split the original large image into multiple sub-images according to a fixed ratio and step size, with a certain overlap between the sub-images; adaptive image splitting is performed by learning the area distribution map of objects appearing at each position on the original image;
[0082] A motion trend judgment module is used to judge the size range of the target based on the heat map of the target's location, the probability of the target appearing at the location, and the target's motion trend;
[0083] A dynamic image segmentation module is used to track the dynamic image of the small target, determine the movement trend of the small target, and segment the dynamic image of the small target according to the movement trend;
[0084] The repeated area suppression module is used to determine whether there is repeated prediction for multiple prediction boxes where the small target exists; if at least one of the prediction boxes is a subbox of other prediction boxes, the target in the overlapping area of the multiple prediction boxes is suppressed according to the area size of the picture where the prediction box where the small target appears is located as a threshold.
[0085] In the self-learning module, during the algorithm initialization self-learning phase, most algorithms currently scale the original large image to a fixed-size small image in the first step. Some small targets (such as ,even ) after scaling ( even ) It is difficult for the naked eye to distinguish, and the algorithm is difficult to recall. In the initialization stage of the early algorithm, in order to avoid the loss of small targets after scaling, the image scaling ratio is reduced. The original image will be divided into several sub-images according to a fixed ratio and step size. At the same time, in order to prevent the objects in the critical area from breaking, each sub-image has a fixed ratio of overlap, and as many small target objects in the original image as possible are recalled, and the area distribution map of the objects appearing at each position on the original image is self-learned.
[0086] The motion trend judgment module judges the size range of the target according to the heat map of the target location, the probability of the target appearing at the location, and the motion trend of the target; judging the size range of the target according to the heat map of the target location, the probability of the target appearing at the location, and the motion trend of the target specifically includes: judging the size range of the target according to the following formula,
[0087]
[0088] in, is the target size evaluation value calculated according to the given conditions, The heat value at the position (x, y) obtained according to the heat map accumulated in the first step; is the absolute value of the ratio of the probability values of the target at position (x, y); where, represents the target probability value at position (x, y), Represents the target probability value in the entire image; Parameters used for adjustment, indicating the size of the slice or the number of cropped sub-pictures; is the height change of the target at position (x, y); is the probability value of the target at position (x, y), is the average entropy value of the heat map; It is the average value of the evaluation value obtained according to the target area change statistics of the tracking module. In step 2 of this embodiment, the target probability ratio reflects the relative probability size of the target at the position by calculating the absolute value of the ratio of the target probability value at the position (x, y) to the target probability value in the entire image. A larger ratio indicates that the probability of the target at this position is higher. The height change value is the height change value of the target at the position (x, y), taking into account the size change of the target in the vertical direction. A larger height change value indicates that the vertical size of the target at this position is larger. The magnitude and range of the target size evaluation value are adjusted by adjusting the parameters and the average entropy value of the heat map.
[0089] A dynamic image segmentation module is used to track the dynamic image of the small target, determine the movement trend of the small target, and segment the dynamic image of the small target according to the movement trend;
[0090] Track the dynamic image of the small target, determine the movement trend of the small target, and segment the dynamic image of the small target according to the movement trend; in step 3, first obtain the original image sequence, which includes continuous frames of multiple time steps, and use the target tracking algorithm, such as Kalman filter, correlation filter or deep learning tracker, to track the small target in the initial frame. The tracker will output the position information of the small target at each time step; according to the output of the tracker, establish the tracking chain of the small target and record the position information of the small target at each time step.
[0091] Determining the movement trend of the small target includes: establishing a complete tracking chain for each small target entering and leaving the field of view, and during the tracking process, monitoring the area change of each object is achieved through the following formula:
[0092]
[0093] in, Indicates the maximum area within the time range of the tracking chain, represents the minimum value of the area within the time range in the tracking chain, β represents the weight factor, represents the absolute value of the change in size or area between adjacent time steps; t represents the index of the time step, from 1 to N; in this formula, the motion trend score of the small target is obtained by calculating the sum of the absolute values of the area change of the small target between adjacent time steps and the difference between the maximum and minimum values of the area of the small target within the observation time range. , the rating It can be used to measure the change in size or area of a small target within the observation time range, so as to determine its movement trend; if the movement trend of the small target is from a small area to a large area, then reduce the number of cut sub-images and increase the sub-image size; if the movement trend of the small target is from a large area to a small area, then increase the number of cut sub-images and reduce the sub-image size. For example, suppose we have a small target tracking chain containing 5 time steps (t1, t2, t3, t4, t5), and we have recorded the area of the small target s(t) at each time step. Calculate the absolute value of the area change between each adjacent time step, where t is 1 to 4. Substituting into the formula, if If the size of the small target is larger, it means that the area of the small target has a significant trend of increasing within the observation time range. It can be inferred that the small target may have entered from a distance and gradually increased. In this case, you can consider reducing the number of cropped sub-images and increasing the sub-image size to better capture the details and features of the small target.
[0094] if If the size of the small target is smaller, it means that the area of the small target changes little or tends to be stable within the observation time range. It can be inferred that the small target may be close and remain relatively stable. In this case, you can consider increasing the number of cropped sub-images and reducing the sub-image size to better capture the shape and motion trajectory of the small target.
[0095] A repeated region suppression module, used for determining whether there is repeated prediction for multiple prediction boxes of the small target;
[0096] At this time, there are three cases. The first case is that in the repeated area suppression module, if at least one of the prediction boxes is a sub-box of other prediction boxes, the area size of the picture where the prediction box where the small target appears is located is used as a threshold, and the targets in the overlapping areas of multiple prediction boxes are suppressed, specifically including: according to the area of a preset threshold, sorting multiple prediction boxes with an area larger than the preset threshold from large to small according to the area; selecting the box with the largest area as the candidate box; calculating the degree of overlap between the candidate box and the multiple prediction boxes, if the degree of overlap between a certain prediction box and the candidate box is greater than the preset threshold, then deleting the boundary box, otherwise retaining it; continuing to execute until all prediction boxes are processed.
[0097] In the second case, in the repeated region suppression module, if there are at least two prediction boxes that are complete boxes including the small target, all prediction boxes are sorted in descending order according to the confidence, and the box with the highest confidence is selected as the candidate box. The degree of overlap between the candidate box and the multiple prediction boxes is calculated. If the degree of overlap between a prediction box and the candidate box is greater than a preset threshold, the bounding box is deleted, otherwise it is retained; continue to execute until all prediction boxes are processed. We have the following three boxes and their confidences:
[0098] Box 1: (x1, y1, x2, y2) = (50, 50, 100, 100), with a confidence level of 0.9
[0099] Box 2: (x1, y1, x2, y2) = (60, 60, 120, 120), with a confidence level of 0.8
[0100] Box 3: (x1, y1, x2, y2) = (70, 70, 90, 90), with a confidence level of 0.7
[0101] Use the non-maximum suppression (NMS) algorithm with confidence as the criterion, for example:
[0102] Sort all boxes in descending order by confidence:
[0103] Box 1: Confidence level 0.9
[0104] Box 2: Confidence level 0.8
[0105] Box 3: Confidence level 0.7
[0106] Select the box with the highest confidence (box 1) and add it to the final list of selected boxes.
[0107] Calculate the intersection over union (IoU) of box 2 and box 1:
[0108] IoU(box 2, box 1) = calculate the intersection area of the two boxes / calculate the union area of the two boxes
[0109] If IoU(box 2, box 1) is greater than the set threshold (for example, 0.5), box 2 is discarded; otherwise, box 2 is added to the final selected box list.
[0110] Calculate the intersection over union (IoU) of box 3 and box 1:
[0111] IoU(box 3, box 1) = calculate the intersection area of the two boxes / calculate the union area of the two boxes
[0112] If IoU(box 3, box 1) is greater than the set threshold (e.g. 0.5), box 3 is discarded; otherwise, box 3 is added to the final selected box list.
[0113] Regarding how to determine the degree of overlap between the candidate box and the multiple prediction boxes, in this embodiment, when calculating the degree of overlap between the candidate box and the multiple prediction boxes by using the intersection-over-union ratio, the intersection-over-union ratio loss function is used. for:
[0114]
[0115] in, is the intersection-over-union ratio of the prediction box; is the target value of the intersection-union ratio.
[0116] The formula is based on the IOU of the predicted bounding box and the target value The comparison is divided into two cases for calculation. If the IOU of the predicted bounding box is less than the target value , then use the first formula to calculate IOU Loss. In this case, is the IOU of the predicted bounding box and the target value If the IOU of the predicted bounding box is greater than or equal to the target value , then use the second formula to calculate .in this case, is the target value The negative of the natural logarithm of the IOU with the predicted bounding box.
[0117] The intersection-over-union loss is used to measure the difference between the predicted bounding box and the target bounding box. When the consistency is good, the intersection-over-union loss approaches 0, indicating that the predicted bounding box has a high degree of overlap with the target bounding box. When the difference between is large, the intersection-over-union loss increases, indicating that the overlap between the predicted bounding box and the target bounding box is low.
[0118] The intersection-over-union loss function can be used to train the target detection model. In the target detection task, the model needs to learn the correspondence between the predicted box and the true labeled box. By using the intersection-over-union loss function, the overlap between the predicted box and the true labeled box can be measured, thereby guiding the model to accurately locate the target and train the bounding box regression.
[0119] In addition, the intersection-over-union loss function can be used as one of the indicators for evaluating the performance of the target detection model. The intersection-over-union ratio between the predicted box and the true labeled box is calculated to measure the accuracy and recall of the model's prediction. Generally, a higher intersection-over-union ratio indicates that the model has better positioning and detection capabilities.
[0120] In this embodiment, when calculating the overlap between the candidate box and the multiple prediction boxes, the intersection-over-union score can be enhanced. Calculate the overlap between the candidate box and the multiple prediction boxes, and the enhanced intersection-over-union score Specifically include:
[0121]
[0122] The formula consists of three terms. The first term measures the distance between the center of the predicted box and the center of the candidate box. , relative to the size of the selection box, Represents the width of the selection box. Represents the height of the candidate box; the second term measures the difference in width between the predicted box and the candidate box, relative to the square of the width of the candidate box; Represents the difference in width between the predicted box and the candidate box; the third term measures the difference in height between the predicted box and the candidate box, relative to the square of the height of the candidate box; Represents the height difference between the predicted box and the candidate box.
[0123] By combining these components, the enhanced intersection-over-union metric can comprehensively consider the position, size, and shape information of the bounding box, providing a more accurate bounding box matching metric. It can be used as a metric for model performance evaluation in object detection tasks, or as a loss function in bounding box regression tasks to help the model learn accurate bounding box predictions.
[0124] There is also a third situation, if the small target overlaps in the adjacent multiple prediction frames, the upper and lower detection frames or the left and right detection frames are merged according to the overlap ratio. In this way, when the large image is cut into sub-images, although there is a certain overlap in the upper and lower and left and right ratios (such as 20% width and height of the sub-image), there are still objects that will be cut off from the upper and lower and left and right. At this time, by calculating the horizontal overlap ratio of the left and right sub-images, the left and right are merged when it is higher than the threshold, and the upper and lower sub-images are also merged in the height direction.
[0125] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
[0126] It is obvious to those skilled in the art that the embodiments of the present invention are not limited to the details of the above exemplary embodiments, and that the embodiments of the present invention can be implemented in other specific forms without departing from the spirit or basic features of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the embodiments of the present invention is limited by the attached claims rather than the above description, so it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the embodiments of the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units, modules or devices stated in the system, device or terminal claims can also be implemented by the same unit, module or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0127] Finally, it should be noted that the above implementation modes are only used to illustrate the technical solutions of the embodiments of the present invention and are not intended to limit them. Although the embodiments of the present invention have been described in detail with reference to the above preferred implementation modes, those skilled in the art should understand that the technical solutions of the embodiments of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A small target detection method on the terminal based on tracking and adaptive image slicing, characterized in that: include: Step 1: Divide the original large image into multiple sub-images according to a fixed ratio and step size, with a certain overlap between the sub-images; Adaptive image cutting is performed by learning the area distribution map of objects appearing at each position on the original image; Step 2, judging the size range of the target based on the heat map of the target's location, the probability of the target appearing at the location, and the target's movement trend; Step 3, tracking the dynamic image of the small target, determining the movement trend of the small target, and segmenting the dynamic image of the small target according to the movement trend; Step 4: for the multiple prediction boxes where the small target exists, determine whether there is repeated prediction; if at least one of the prediction boxes is a subbox of other prediction boxes, use the area size of the picture where the prediction box where the small target appears is located as a threshold to suppress the targets in the overlapping areas of the multiple prediction boxes; Among them, in the step 4, if at least one of the prediction boxes is a sub-box of other prediction boxes, according to the area size of the picture where the prediction box where the small target appears is located as a threshold, suppressing the target in the overlapping area of multiple prediction boxes specifically includes: according to the area of a preset threshold, sorting multiple prediction boxes with an area larger than the preset threshold from large to small according to the area; selecting the box with the largest area as the candidate box; calculating the degree of overlap between the candidate box and the multiple prediction boxes, if the degree of overlap between a certain prediction box and the candidate box is greater than the preset threshold, deleting the boundary box, otherwise retaining it; continuing to execute until all prediction boxes are processed; If there are at least two prediction boxes that are complete boxes including the small target, all prediction boxes are sorted in descending order according to the confidence, the box with the highest confidence is selected as the candidate box, and the degree of overlap between the candidate box and the multiple prediction boxes is calculated. If the degree of overlap between a prediction box and the candidate box is greater than a preset threshold, the bounding box is deleted, otherwise it is retained; continue to execute until all prediction boxes are processed; If there is overlap of the small objects in multiple adjacent prediction frames, the upper and lower detection frames or the left and right detection frames are merged according to the overlap ratio.
2. The on-device small target detection method based on tracking and adaptive graph slicing according to claim 1, characterized in that: In step 2, judging the size range of the target according to the heat map of the target location, the probability of the target appearing at the location, and the movement trend of the target specifically includes: judging the size range of the target according to the following formula: Among them, g′ is the target size evaluation value calculated according to the given conditions, and m′ is the thermal value at the position (x, y) obtained according to the thermal map accumulated in the first step; is the absolute value of the ratio of the probability values of the target at the position (x, y); wherein v′ represents the target probability value at the position (x, y), and v represents the target probability value in the entire image; l represents the parameter used for adjustment, indicating the size of the cut image or the number of cropped sub-images; Δh′ is the height change value of the target at the position (x, y); p′ is the probability value of the target at the position (x, y), and e is the average entropy value of the heat map; n′ is the average value of the evaluation value obtained according to the target area change statistics of the tracking module.
3. The on-device small target detection method based on tracking and adaptive graph slicing according to claim 1, characterized in that: Determining the movement trend of the small target in step 3 includes: establishing a complete tracking chain for each small target entering and leaving the field of view, and during the tracking process, monitoring the area change of each object is achieved through the following formula: Among them, max(s(t)) represents the maximum area within the time range of the tracking chain, min(s(t)) represents the minimum area within the time range of the tracking chain, β represents the weight factor, |s(t+1)-s(t)| represents the absolute value of the change in size or area between adjacent time steps; t represents the index of the time step, from 1 to N; In this formula, the motion trend score S of the small target is obtained by calculating the sum of the absolute values of the area change of the small target between adjacent time steps and the difference between the maximum and minimum values of the area of the small target within the observation time range. s , the score S s It can be used to measure the change in size or area of a small target within the observation time range, so as to determine its movement trend; if the movement trend of the small target is from a small area to a large area, the number of cut sub-images is reduced and the sub-image size is increased; if the movement trend of the small target is from a large area to a small area, the number of cut sub-images is increased and the sub-image size is reduced.
4. The on-device small target detection method based on tracking and adaptive graph cutting according to claim 1, characterized in that: When calculating the overlap between the candidate box and the multiple prediction boxes by using the intersection-over-union ratio, the intersection-over-union ratio loss function R is used. IOU loss is applied in the model, in the intersection-over-union loss function: Among them, P IOU is the intersection-over-union ratio of the prediction box; IOU tar is the target value of the intersection-union ratio.
5. The on-device small target detection method based on tracking and adaptive graph slicing according to claim 1, characterized in that: By enhancing the intersection-union ratio score R EIoU Calculate the overlap between the candidate box and the multiple prediction boxes, and the enhanced intersection-over-union score R EIoU Specifically include: The formula consists of three terms. The first term measures the distance D between the center of the predicted box and the center of the candidate box. centers , relative to the size of the selection box, W c Represents the width of the selection box, H c represents the height of the candidate box; the second term measures the difference in width between the predicted box and the candidate box, relative to the square of the width of the candidate box; D w represents the difference in width between the predicted box and the candidate box; the third item measures the difference in height between the predicted box and the candidate box, relative to the square of the height of the candidate box; D h Represents the height difference between the predicted box and the candidate box.
6. A small target detection device on the terminal based on tracking and adaptive image slicing, characterized in that: include: The self-learning module is used to split the original large image into multiple sub-images according to a fixed ratio and step size, and there is a certain overlap between the sub-images; Adaptive image cutting is performed by learning the area distribution map of objects appearing at each position on the original image; A motion trend judgment module is used to judge the size range of the target based on the heat map of the target's location, the probability of the target appearing at the location, and the target's motion trend; A dynamic image segmentation module is used to track the dynamic image of the small target, determine the movement trend of the small target, and segment the dynamic image of the small target according to the movement trend; A repeated area suppression module is used to determine whether there is repeated prediction for multiple prediction boxes where the small target exists; if at least one of the prediction boxes is a subbox of other prediction boxes, the target in the overlapping area of the multiple prediction boxes is suppressed according to the area size of the picture where the prediction box where the small target appears is located as a threshold; The repeated area suppression module is further used for: if at least one of the prediction boxes is a sub-box of other prediction boxes, suppressing the target in the overlapping area of multiple prediction boxes according to the area size of the picture where the prediction box where the small target appears is located as a threshold, specifically including: sorting multiple prediction boxes with an area larger than the preset threshold according to the area from large to small according to the area; selecting the box with the largest area as the candidate box; calculating the degree of overlap between the candidate box and the multiple prediction boxes, if the degree of overlap between a certain prediction box and the candidate box is greater than the preset threshold, deleting the boundary box, otherwise retaining it; continuing to execute until all prediction boxes are processed; If there are at least two prediction boxes that are complete boxes including the small target, all prediction boxes are sorted in descending order according to the confidence, the box with the highest confidence is selected as the candidate box, and the degree of overlap between the candidate box and the multiple prediction boxes is calculated. If the degree of overlap between a prediction box and the candidate box is greater than a preset threshold, the bounding box is deleted, otherwise it is retained; continue to execute until all prediction boxes are processed; If there is overlap of the small objects in multiple adjacent prediction frames, the upper and lower detection frames or the left and right detection frames are merged according to the overlap ratio.
Citation Information
Patent Citations
Pin-level defect identification method of unmanned aerial vehicle inspection image
CN115393264A
Global perception small target intelligent detection method for public safety video
CN115410060A