Process method for detecting tiny object
By monitoring small object losses in real time and triggering adaptive strategies, combining the feature fusion of pre-trained large models and dynamic adjustment of reinforcement learning agents, the problem of insufficient accuracy and robustness of small-object detection in the existing technology is solved, and a small-object detection method with high accuracy and continuous optimization is achieved.
Patent Information
- Application Number
- CN202510222108.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
It is difficult for existing small object detection technology to pay attention to and adaptively regulate the loss of small objects in real time, resulting in high missed detection rate, low confidence and difficult to improve detection accuracy.
By monitoring the loss of small objects in real time, data augmentation and loss-weighted adaptive strategies are automatically triggered, combined with the global features of pre-trained large models to fuse with multi-scale feature pyramids, and dynamically adjust data augmentation parameters using reinforcement learning agents.
It significantly improves the accuracy and robustness of small object detection, reduces false detection and missed detection of small objects, and realizes continuous online update and optimization of detection methods through cloud-based large-scale model services.
Smart Images

Figure CN120147722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual recognition, and specifically to a process method for detecting tiny objects. Background Art
[0002] Existing small target detection technologies mainly follow the general target detection framework. In data augmentation, only simple random cropping or scaling operations are performed. At the same time, multi-scale feature extraction is introduced in the network structure, but little attention is paid to small objects in real time and adaptive regulation is not carried out.
[0003] Specifically, these methods generally do not separately monitor the loss situation of small objects during the classification and positioning processes, nor do they set additional loss weighting or cross-attention mechanisms for such targets. Even in the post-processing stage, usually only a unified non-maximum suppression strategy is adopted, and the scale difference between small objects and large objects cannot be effectively distinguished. Due to the lack of multi-level optimization means such as specialized reinforcement learning or cloud collaboration, when facing situations such as long-distance shooting, complex scenes, or dense targets, problems such as high false negative rates, low confidence levels, and difficulty in improving detection accuracy often occur, and the overall performance is not satisfactory.
[0004] The difficulties in small target detection mainly lie in the small size of the target, weak feature expression ability, and being easily interfered by the background and noise. In addition, the scale of small targets varies greatly, key information is easily lost, and features are easily lost during the downsampling process. Current research mainly focuses on improving feature extraction capabilities (such as multi-scale feature fusion, feature enhancement), improving loss functions (such as focal loss), and optimizing detection box strategies (such as dense anchor boxes or Transformer-based methods) to improve detection accuracy and robustness. However, current methods all have certain limitations in detecting small target objects, so an intelligent algorithm is needed to detect tiny objects. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the existing technology, the present invention provides a process method for detecting tiny objects, which automatically triggers an adaptive strategy of data augmentation and loss weighting by monitoring the loss situation of small objects in real time. At the same time, a pre-trained large model (such as ViT / CLIP) is introduced, and its global features are fused with the multi-scale feature pyramid to further improve the recognition ability of small objects. We also use a reinforcement learning agent to dynamically adjust the data augmentation parameters according to the actual situation to make the entire detection process more intelligent.
[0007] (2) Technical Solutions
[0008] To achieve the above object, the present invention provides the following technical solution: A process method for detecting tiny objects, comprising the following steps:
[0009] Real-time monitor the loss of small objects, dynamically adjust the attention to small objects by combining the set threshold, select multiple images including small objects from the training set for splicing, and obtain a multi-image splicing sample;
[0010] In the detection head, add a learnable or dynamically adjustable weight coefficient to the loss term for small objects, and combine the monitored loss value of small objects to increase or decrease the weight coefficient, and assist in feature extraction based on the pre-trained large model;
[0011] Construct a multi-scale feature pyramid, integrate the global information of the pre-trained large model, monitor the detection effects of each scale in real time, and perform dynamic adjustment strategies;
[0012] Through the reinforcement learning agent, perform adaptive data augmentation parameter tuning and achieve online adaptive update;
[0013] Perform non-maximum suppression on the detection results, perform confidence calibration, and perform iterative feedback and online fine-tuning;
[0014] Deploy the large model to the cloud, provide real-time global feedback, and continuously update online.
[0015] Furthermore, real-time monitor the loss of small objects, and set a threshold to dynamically adjust the attention to small objects, including:
[0016] After each training iteration or the end of a sample set, extract the loss values corresponding to small objects from the classification loss and the localization loss respectively;
[0017] According to pre-statistics or experiments, set one or a group of dynamic thresholds. When the classification and localization losses of small objects are continuously lower than the threshold, it is regarded as insufficient attention to small objects;
[0018] In multiple training iterations, if it is continuously observed that the loss value of small objects is lower than the threshold, it is determined that more attention needs to be paid to small objects.
[0019] Furthermore, select multiple images including small objects from the training set for splicing to obtain a multi-image splicing sample, including:
[0020] Splice multiple images onto the same blank canvas, and place them in the new image according to the positions of the small objects in the original images;
[0021] Increase the proportion of small objects, merge multiple images into one picture, and increase the number of small objects in the picture;
[0022] Image scaling and restoration, scaling to the original size: After splicing is completed, scale the merged image again;
[0023] Simulate the telephoto / fine-grained scenario. Since small objects before stitching may have been enlarged after merging, shrinking them back to the standard size can simulate the visual characteristics in the telephoto or small-target scenario, enabling the model to gradually adapt to the detection of small objects in complex scenarios.
[0024] Furthermore, introduce a random scaling factor when scaling the image.
[0025] Furthermore, in combination with the monitored loss value of small objects, increase or decrease the weight coefficient, including:
[0026] When the monitored loss value of small objects continues to be low, automatically increase the weight coefficient so that the gradient from small objects is larger during backpropagation;
[0027] When the monitored loss value of small objects keeps rising, decrease the weight.
[0028] Furthermore, based on the pre-trained large model, assist in feature extraction, including:
[0029] Introduce a pre-trained large model, select Vision Transformer, CLIP or other models trained on large-scale image data, fix the backbone parameters of the large model, and only perform a small amount of fine-tuning on the detection task data, such as adapting to specific channel numbers and resolutions;
[0030] Feature extraction and fusion to obtain global semantics. Forward-propagate the input image in the large model to obtain a high-level global semantic vector or feature map, and perform cross-attention fusion by introducing a cross-attention module at a certain layer such as the detection backbone or the FPN layer;
[0031] The query (Q) can come from the intermediate features of the detection backbone, and the key (K) and value (V) come from the global features of the large model;
[0032] Through attention calculation, inject the global information extracted by the large model into the feature representation of the local detection network;
[0033] Maintain fine-grained features. The fused features contain both convolutional features with relatively high local resolution and the global context of the large model.
[0034] Furthermore, build a multi-scale feature pyramid, integrate the global information of the pre-trained large model, monitor the detection effects of each scale in real time, and perform dynamic adjustment strategies, including:
[0035] Output feature maps with different resolutions at different stages of common detection backbone networks, build a pyramid through top-down sampling and lateral connections to ensure the scale difference between levels, improve the resolution of the bottom feature map or introduce Deformable convolution, and perform classification and regression respectively at different P layers through shared or separate detection heads;
[0036] First, perform "re - mapping" on the global semantic vectors or multi - scale features extracted from the large model to align with the number of channels or spatial resolutions of different scales of the FPN. Adopt an adaptive attention mechanism to weight each scale feature of the FPN through the learned weights. Design an MLP or lightweight network and input the global features output by the large model into it. The MLP outputs a set of weight vectors (α1, α2, α3, α4, α5), corresponding to the five scales of the FPN. During feature fusion, the output of the corresponding layer can be expressed as: P′i = αi×Pi or P′i = αi×Pi + βi×GlobalFeat;
[0037] During training or validation, statistically calculate the detection Recall, Precision, or mAP of small objects for each FPN scale. If it is found that the detection effect of a certain scale on small objects is significantly insufficient, it is necessary to increase the weight of this scale in the fusion. According to the monitoring results, perform a secondary correction on αi: αi←αi + Δαi.
[0038] Furthermore, through a reinforcement learning agent, perform adaptive data augmentation parameter tuning and achieve online adaptive updates, including:
[0039] Set the current detection performance metrics of the model for small objects, the current data augmentation parameters, adjust the hyperparameter combination of data augmentation, and give positive or negative rewards;
[0040] Embed the RL agent into the training loop, call it periodically or at a certain number of steps. The RL agent calculates the Reward based on the current detection results, updates the policy, and outputs the data augmentation parameters for the next stage. If not only small objects are concerned but also large objects or the overall mAP, a weighted reward function can be introduced;
[0041] The RL agent may perform more random explorations in the early stage of training; as the number of epochs increases, the agent will form a stable preference for effective augmentation strategies, and transmit the RL decision results to the data loading and augmentation pipeline in real - time. The performance of the trained model in turn updates the RL state, realizing a closed - loop feedback.
[0042] Furthermore, perform non - maximum suppression on the detection results, perform confidence calibration, and perform iterative feedback and online fine - tuning, including:
[0043] Before performing NMS on the detection results, first group them according to the size of the predicted bounding boxes. Set a relatively smaller IoU threshold for the group of small objects, or use methods such as Soft - NMS, DIoU - NMS, etc., to better retain adjacent small targets. For the group of large objects, maintain the conventional NMS threshold. If small objects are dense in the detection scene, use methods such as weighted box fusion to replace traditional hard suppression, and perform intelligent fusion according to the confidence and position deviation output by the detection head;
[0044] The feature ROI of the candidate box or the corresponding FPN feature is fused with the global semantic vector extracted by the large model again. Through a lightweight MLP or a weighting function, the confidence is corrected twice. For small objects lacking local features but strongly related to the global context, their confidence is moderately increased. According to the accuracy performance of small and large objects on the validation set, the classification scoring threshold is dynamically changed, the threshold for small objects is appropriately reduced, and the threshold for large objects is maintained or increased;
[0045] The missed detection or misdetection samples of small objects in the detection stage are re-annotated or automatically screened, and the dataset is updated in an online / semi-online manner to enrich the small object sample distribution. In the deployment stage, fine-tuning is performed regularly according to the real-time detection results. During fine-tuning, most of the weights of the backbone and the large model are kept, and mainly the detection head and the weights related to small objects are updated to quickly iterate and optimize the small object detection ability.
[0046] Furthermore, the large model is deployed to the cloud for real-time global feedback and continuous online update, including:
[0047] The cloud server can continuously collect data from different application scenarios and different geographical regions to perform periodic retraining or incremental training on the large model. The local system only needs to regularly upload necessary feature statistics or desensitized detection samples and performance metrics to the cloud. The cloud returns the updated pre-trained weights or feature fusion strategies according to the global information;
[0048] The local detection system interacts with the cloud after a set period or when certain triggering conditions are met, pulls the latest global feature extraction layer weights, large model parameters, and data augmentation strategy suggestions. According to the update package sent by the cloud, the local system performs real-time fine-tuning or replaces some modules to form an "cloud-edge" collaborative online evolution mechanism;
[0049] When the local detection scenario changes, the cloud can quickly adapt and generate new parameters after receiving new data.
[0050] (III) Beneficial Effects
[0051] Compared with the prior art, the present invention provides a process method for detecting tiny objects, having the following beneficial effects:
[0052] 1. During the detection process, by real-time monitoring the loss situation of small objects, an adaptive strategy of data augmentation and loss weighting is automatically triggered. At the same time, a pre-trained large model (such as ViT / CLIP) is introduced, and its global features are fused with the multi-scale feature pyramid to further improve the recognition ability of small objects. We also use a reinforcement learning agent to dynamically adjust the data augmentation parameters according to the actual situation, making the whole detection process more intelligent.
[0053] 2. In the post-detection processing stage, differential NMS and confidence calibration techniques are adopted to effectively reduce false detections and missed detections of small objects. In addition, with the help of cloud-based large model services, continuous online updates and optimizations of the detection method in different scenarios are achieved.
[0054] 3. This method focuses on small object detection. It can not only automatically regulate the detection process but also integrates multiple attention mechanisms, and cooperates with cloud services to achieve collaborative iteration, significantly improving the accuracy and robustness of small object detection. Description of the Drawings
[0055] Figure 1 It is a schematic diagram of a process method for detecting tiny objects proposed by the present invention. Detailed Embodiments
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] Please refer to Figure 1 , the present invention provides the following technical solutions: A process method for detecting tiny objects, including the following steps:
[0058] S1. Real-time monitor the loss of small objects, dynamically adjust the attention to small objects in combination with the set threshold, select multiple images including small objects from the training set for splicing, and obtain a multi-image splicing sample;
[0059] S2. In the detection head, add a learnable or dynamically adjustable weight coefficient to the loss term of small objects, and increase or decrease the weight coefficient in combination with the monitored loss value of small objects, and assist in feature extraction based on the pre-trained large model;
[0060] S3. Construct a multi-scale feature pyramid, integrate the global information of the pre-trained large model, real-time monitor the detection effects of each scale, and perform dynamic adjustment strategies;
[0061] S4. Through a reinforcement learning agent, perform adaptive data augmentation parameter tuning and achieve online adaptive updates;
[0062] S5. Perform non-maximum suppression on the detection results, perform confidence calibration, and perform iterative feedback and online fine-tuning;
[0063] S6. Deploy the large model to the cloud, perform real-time global feedback, and continuously update online.
[0064] In S1, the loss of small objects is monitored in real time, and a threshold is set to dynamically adjust the attention to small objects, specifically including:
[0065] 1.1. Detection of Loss Threshold Judgment:
[0066] Monitoring and sampling of small object Loss;
[0067] Definition of small objects: Usually, small, medium, and large objects can be distinguished according to the pixel area of the target bounding box or the ratio relative to the image size.
[0068] Real-time monitoring of small object Loss: After each training iteration or the end of a batch, the loss values corresponding to small objects are respectively extracted from the classification loss (such as cross-entropy or focal loss) and the localization loss (such as L1, Smooth L1, or IoU-based loss).
[0069] Setting the threshold: According to pre-statistics or experiments, set one or a group of dynamic thresholds. When both the classification and localization losses of small objects are continuously lower than the threshold, it is regarded as insufficient attention to small objects.
[0070] Conditions for triggering dynamic adjustment:
[0071] In multiple training iterations, if it is continuously observed that the small object Loss is lower than the threshold, it is determined that more attention needs to be paid to small objects.
[0072] Smoothing means such as EMA (Exponential Moving Average) can be combined to avoid overly frequent adjustments caused by abnormal Loss value fluctuations once or several times.
[0073] Select multiple images including small objects from the training set for splicing to obtain a multi-image splicing sample, including:
[0074] 1.2. Data Input Adjustment:
[0075] Selection of multi-image splicing samples: Select images containing small objects from the training set. When splicing, preferentially select samples with a higher frequency of small object appearance or a rich number of annotations.
[0076] Splicing rules:
[0077] Splice 4 images onto the same blank canvas, and reasonably place them in the new image according to the position of the small objects in the original images, ensuring that the relative distribution of small objects in the new image is denser;
[0078] Increase the proportion of small objects: By combining 4 images into one picture, we can increase the number of small objects in the picture, thereby increasing the chance for the network to detect and recognize small objects in a single forward pass.
[0079] Image scaling and restoration:
[0080] Scaling to the original size: After stitching is completed, the merged image is scaled again (for example, scaled proportionally to the original input size, such as 640×640 or 1024×1024, etc.);
[0081] Simulating telephoto / fine-grained scenarios: Since small objects before stitching may have been enlarged after merging, shrinking them back to the standard size can simulate the visual characteristics in telephoto or small target scenarios, enabling the model to gradually adapt to the detection of small objects in complex scenarios;
[0082] Ensuring diversity: A random scaling factor (such as ranging from 0.8 to 1.2) can be introduced during scaling to further enhance data diversity.
[0083] In S2, adaptive loss weighting and large model feature extraction are performed. Among them, in combination with the detected loss value of small objects, the weight coefficient is increased or decreased, including:
[0084] 2.1 Dynamic loss weighting mechanism:
[0085] Calculation of the weighting factor;
[0086] In the detection head (or the detection branch of the FPN layer), a learnable or dynamically adjustable weight coefficient w_small is added to the loss term for small objects;
[0087] When it is monitored that the small object Loss remains low, w_small is automatically increased, making the gradient from small objects larger during backpropagation;
[0088] Conversely, if it is subsequently found that the small object Loss keeps rising, the weight can be slightly reduced to prevent overfitting or unstable training caused by too large a gradient.
[0089] Specific implementation:
[0090] It can be written in the loss function as:
[0091] Loss = w_small × Loss_small + Loss_normal;
[0092] Among them, Loss_small only counts the detection frames labeled as small objects, and Loss_normal is the conventional overall loss;
[0093] w_small can be dynamically adjusted through threshold triggering + exponential decay or PID (Proportional-Integral-Derivative) control strategy.
[0094] Based on the pre-trained large model, assist in feature extraction, including:
[0095] 2.2. Feature extraction assisted by the large model:
[0096] Introduce the pre-trained large model;
[0097] It is possible to select Vision Transformer (ViT), CLIP or other models trained on large-scale image data;
[0098] The backbone parameters of the large model can be fixed, and only a small amount of fine-tuning is performed on the detection task data (such as adapting to specific channel numbers and resolutions), saving training costs.
[0099] Feature extraction and fusion;
[0100] Obtain global semantics: Forward propagate the input image in the large model to obtain high-level global semantic vectors or feature maps;
[0101] Cross-Attention fusion:
[0102] Introduce a cross-attention module at a certain layer of the detection backbone (such as ResNet, Darknet, Swin Transformer, etc.) or the FPN layer;
[0103] The query (Q) can come from the intermediate features of the detection backbone, and the key (K) and value (V) come from the global features of the large model.
[0104] Through attention calculation, inject the global information extracted by the large model into the feature representation of the local detection network.
[0105] Maintain fine-grained features: The fused features contain both convolutional features with relatively high local resolution and the global context of the large model.
[0106] S3 can achieve multi-scale feature fusion and global context integration. Among them, construct a multi-scale feature pyramid, integrate the global information of the pre-trained large model, monitor the detection effects of each scale in real time, and perform dynamic adjustment strategies, including:
[0107] 3.1. Construction of the multi-scale feature pyramid:
[0108] FPN structure:
[0109] Output feature maps with different resolutions (such as C1, C2, C3, C4, C5) at different stages of common detection backbone networks (such as ResNet50, ResNeXt, Swin Transformer, etc.);
[0110] Construct a pyramid (P1, P2, P3, P4, P5) through top-down sampling and lateral connections to ensure the scale difference between levels.
[0111] Goal: Let the lower-level P layers focus on tiny features (small objects), and the higher-level P layers focus on semantic features (large objects and global semantics).
[0112] Ensure the fine-grainedness of small objects;
[0113] Appropriately increase the resolution of the bottom-level feature maps or introduce improvements such as Deformable Convolution to enhance the network's ability to depict the edge and shape information of small objects.
[0114] Through shared or separate detection heads, perform classification and regression on different P layers respectively to improve the perception of objects of different sizes.
[0115] 3.2. Global information integration of large models:
[0116] Fusion idea;
[0117] First, perform "re-mapping" on the global semantic vectors or multi-scale features extracted from the large model to align with the number of channels or spatial resolution of different scales of FPN;
[0118] Adopt an adaptive attention mechanism to weight each scale feature of FPN with the learned weights.
[0119] Generation of weighting factors;
[0120] Design an MLP or lightweight network and input the global features output by the large model into it.
[0121] The MLP outputs a set of weight vectors (α1, α2, α3, α4, α5), corresponding to the five scales of FPN.
[0122] During feature fusion, the output of the corresponding layer can be expressed as:
[0123] P′i = αi × Pi or P′i = αi × Pi + βi × GlobalFeat (βi can also be learned or set uniformly).
[0124] 3.3. Scale feedback mechanism:
[0125] Monitor the detection effects of each scale in real time;
[0126] During training or validation, count the detection Recall, Precision, or mAP of small objects for each FPN scale.
[0127] If it is found that the detection effect of a certain scale (such as the bottom-level P2) on small objects is significantly insufficient, the weight of this scale in the fusion needs to be increased.
[0128] Dynamic adjustment strategy;
[0129] According to the monitoring results, perform a secondary correction on αi:
[0130] αi ← αi + Δαi
[0131] Or in the next training round, increase the structural parameters such as the number of channels and receptive field of this scale;
[0132] In the inference stage, the confidence of different detection branches can also be allocated according to the target size to ensure that small objects rely more on the low-level high-resolution features.
[0133] The adaptive data augmentation strategy based on reinforcement learning in S4, where, through the reinforcement learning agent, the adaptive data augmentation parameter tuning is performed and the online adaptive update is realized, including:
[0134] 4.1. Design of the reinforcement learning agent:
[0135] State;
[0136] The current detection performance metrics of the model for small objects (such as small object mAP, Recall, localization accuracy, etc.);
[0137] The current data augmentation parameter settings (such as the number of image mosaics, mosaic method, scaling ratio, cropping range, random noise intensity, etc.).
[0138] Action;
[0139] Adjust the hyperparameter combination of data augmentation, such as Mosaic mosaic method, Mixup transparency, random cropping ratio, etc.;
[0140] It can be discretized (such as several alternative solutions) or continuousized (using the policy network to output specific numerical values).
[0141] Reward;
[0142] Positive reward: If the detection accuracy of small objects improves after this round of training, give a positive reward to the reinforcement learning agent;
[0143] Negative reward: If the accuracy drops or overfitting occurs, give a negative reward to limit adverse adjustments.
[0144] 4.2. Adaptive data augmentation parameter tuning:
[0145] Embedding in the training process;
[0146] 1. Embed the RL agent into the training loop and call it periodically or at a certain number of steps:
[0147] The RL agent calculates the Reward based on the current detection results;
[0148] The RL agent updates the policy (Policy Gradient, Q-learning, etc.);
[0149] The RL agent outputs the data augmentation parameters for the next stage.
[0150] 2. As the training progresses, the RL agent will gradually converge to a better augmentation policy.
[0151] Multi-objective collaboration:
[0152] If not only small objects are concerned, but also large objects or the overall mAP are of interest, a weighted reward function can be introduced:
[0153] R = ωs × mAP_small + ωm × mAP_medium + ωl × mAP_large;
[0154] This can balance the overall detection performance and the priority of small object detection.
[0155] 4.3. Online adaptive update:
[0156] Continuous trial and error and convergence;
[0157] The RL agent may do more random exploration in the early stage of training; as the number of epochs increases, the agent will form a stable preference for effective augmentation policies;
[0158] Avoid the failure of fixed augmentation policies in the later stage.
[0159] Closed-loop with the data preprocessing module;
[0160] Transmit the RL decision results to the data loading and augmentation pipeline in real time;
[0161] The performance of the trained model in turn updates the RL state, achieving closed-loop feedback.
[0162] Post-processing optimization and online feedback fine-tuning mechanism in S5, where non-maximum suppression is performed on the detection results, confidence calibration is carried out, and iterative feedback and online fine-tuning are included:
[0163] 5.1. Improved non-maximum suppression (NMS):
[0164] Distinguish small objects from large objects;
[0165] Before performing NMS on the detection results, group them first according to the size of the prediction boxes;
[0166] Set a relatively smaller IoU threshold for grouping small objects, or use methods such as Soft-NMS and DIoU-NMS to better retain adjacent small targets;
[0167] For grouping large objects, maintain the conventional NMS threshold to avoid excessive box overlaps.
[0168] Hybrid suppression strategy:
[0169] If small objects are dense in the detection scene, methods such as Weighted Box Fusion can be used to replace traditional hard suppression;
[0170] Perform intelligent fusion based on the confidence and position deviation output by the detection head to further reduce missed detections and improve positioning accuracy.
[0171] 5.2. Confidence calibration mechanism:
[0172] Utilize the global context of the large model;
[0173] Fuse the feature ROI of the candidate box (BBox) or the corresponding FPN feature with the global semantic vector extracted by the large model again;
[0174] Perform a secondary correction on the confidence through a lightweight MLP or a weighted function;
[0175] For small objects lacking local features but strongly related to the global context, moderately increase their confidence to further reduce missed detections.
[0176] Dynamically adjust the confidence threshold;
[0177] The classification scoring threshold (scorethreshold) can be dynamically changed according to the accuracy performance of small and large objects on the validation set;
[0178] Appropriately lower the threshold for small objects to increase the recall rate; maintain or increase the threshold for large objects to prevent false detections.
[0179] 5.3. Iterative feedback and online fine-tuning:
[0180] Collect error samples;
[0181] Perform secondary annotation or automatic screening on the missed detection or misdetection samples of small objects in the detection stage;
[0182] Update the dataset in an online / semi-online manner to enrich the sample distribution of small objects.
[0183] Periodically retrain or fine-tune
[0184] In the deployment stage, fine-tuning can be performed regularly according to the real-time detection results;
[0185] During fine-tuning, most of the weights of the backbone and the large model can be maintained, and mainly the detection head and the weights related to small objects are updated to quickly iterate and optimize the small object detection ability.
[0186] The real-time feedback update mechanism based on the cloud large model service in S6, in which the large model is deployed to the cloud for real-time global feedback and continuous online update, including:
[0187] 6.1. Cloud large model integration:
[0188] Cloud training platform;
[0189] The cloud server can continuously collect data from different application scenarios and different geographical regions;
[0190] Periodically retrain or incrementally train the large model (such as CL IP, ViT, etc.) to ensure its generalization ability for new scenarios and new target types.
[0191] Data interface design;
[0192] The local system only needs to regularly upload necessary feature statistics or desensitized detection samples and performance metrics to the cloud;
[0193] Based on the global information, the cloud returns the updated pre-trained weights or feature fusion strategies (such as attention weight distribution).
[0194] 6.2. Real-time global feedback:
[0195] Obtain the latest optimization suggestions;
[0196] The local detection system can interact with the cloud at a set period (such as daily / weekly) or after meeting certain trigger conditions;
[0197] Pull the latest global feature extraction layer weights, large model parameters, data augmentation strategy suggestions, etc.
[0198] Dynamically adjust the local network
[0199] According to the update package sent by the cloud, the local system can fine-tune or replace some modules in real time (such as the projection layer, MLP, and RL agent initialization weights of the large model);
[0200] Form an online evolution mechanism of "cloud-edge" collaboration.
[0201] 6.3. Continuous online update:
[0202] Cross-scene and cross-device;
[0203] When the local detection scenario changes (such as camera resolution replacement, monitoring scenario migration, etc.), the cloud can quickly adapt and generate new parameters after receiving new data;
[0204] Ensure that devices in various locations can share global knowledge and improve the overall small object detection effect.
[0205] Adaptive and scalable;
[0206] This cloud-local collaboration method allows for the rapid introduction of new algorithms (such as the latest large model structure, attention fusion method) and their distribution within a short period of time;
[0207] Enable the small object detection model to maintain high accuracy and robustness even under limited hardware resources or changing deployment scenarios.
[0208] The beneficial effects of the present invention are as follows: Compared with the existing technology, the present invention conducts targeted adaptive optimization in multiple links for small object detection. At the data layer, by monitoring the loss of small objects in real time, data augmentation and loss weighting are triggered. At the same time, reinforcement learning is used to dynamically adjust the augmentation strategy to make data processing more in line with actual needs. At the model layer, the global features of the large model and the multi-scale feature pyramid are fused to improve the fine-grained recognition ability of small objects. In the post-processing stage, differential NMS and confidence calibration are adopted to reduce false detections and missed detections of small objects.
[0209] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A process method for detecting tiny objects, characterized in that: The following steps are involved: Monitor the loss of small objects in real time, dynamically adjust the attention of small objects based on the set threshold, select multiple images including small objects from the training set for stitching, and obtain multi-image stitching samples; In the detection head, a learnable or dynamically adjustable weight coefficient is added to the loss term of small objects. The weight coefficient is increased or decreased based on the monitored small object loss value to assist feature extraction based on the pre-trained large model. Construct a multi-scale feature pyramid, integrate the global information of the pre-trained large model, monitor the detection effect of each scale in real time, and dynamically adjust the strategy; Through reinforcement learning agents, adaptive data enhancement parameter tuning is performed and online adaptive updates are achieved; Perform non-maximum suppression on the detection results, calibrate the confidence, and perform iterative feedback and online fine-tuning; Deploy large models to the cloud, get real-time global feedback, and continuously update them online.
2. A process method for detecting tiny objects according to claim 1, characterized in that: Monitor small object loss in real time and set thresholds to dynamically adjust small object attention, including: After each training iteration or a sample set, the loss values corresponding to small objects are extracted from the classification loss and the localization loss respectively; According to pre-statistics or experiments, set one or a group of dynamic thresholds. When the classification and positioning losses of small objects are continuously lower than the threshold, it is considered that the attention paid to small objects is insufficient. In multiple training iterations, if the loss value of small objects is continuously observed to be lower than the threshold, it is determined that more attention needs to be paid to small objects.
3. A process method for detecting tiny objects according to claim 1, characterized in that: Select multiple images including small objects from the training set for stitching to obtain multi-image stitching samples, including: Stitch multiple images onto the same blank canvas, and place small objects in the new image based on their positions in the original image; Improve the proportion of small objects, merge multiple images into one picture, and increase the number of small objects in the picture; Image scaling and restoration, scaling to original size: after stitching is completed, scale the merged image again; Simulate telephoto / fine-grained scenes. Since small objects before stitching may have been enlarged after merging, shrinking them back to the standard size can simulate the visual characteristics of telephoto or small target scenes, allowing the model to gradually adapt to small object detection in complex scenes.
4. A process method for detecting tiny objects according to claim 3, characterized in that: Introduces a random scaling factor when scaling an image.
5. A process method for detecting tiny objects according to claim 1, characterized in that: Combined with the monitored small object loss value, increase or decrease the weight coefficient, including: When it is detected that the loss value of small objects is continuously low, the weight coefficient is automatically increased to make the gradient from small objects larger during back propagation; When the small object loss value is detected to be increasing, the weight is reduced.
6. A process method for detecting tiny objects according to claim 5, characterized in that: Based on the pre-trained large model, auxiliary feature extraction, including: Introduce a pre-trained large model, select Vision Transformer, CLIP or other models trained on large-scale image data, fix its backbone parameters through the large model, and only perform a small amount of fine-tuning on the detection task data, such as adapting to a specific number of channels and resolution; Feature extraction and fusion, obtain global semantics, forward propagate the input image in the large model, obtain high-level global semantic vectors or feature maps, cross-attention fusion, and introduce cross-attention modules in a certain layer such as the detection backbone or FPN layer; The query (Q) can come from the intermediate features of the detection backbone, and the key (K) and value (V) come from the global features of the large model; Through attention calculation, the global information extracted by the large model is injected into the feature representation of the local detection network; Keeping fine-grained features, the fused features contain both the convolutional features with higher local resolution and the global context of the large model.
7. A process method for detecting tiny objects according to claim 1, characterized in that: Construct a multi-scale feature pyramid, integrate the global information of the pre-trained large model, monitor the detection effect of each scale in real time, and dynamically adjust the strategy, including: Output feature maps of different resolutions at different stages of the common detection backbone network, build a pyramid through top downsampling and lateral connections to ensure scale differences between levels, improve the resolution of the underlying feature maps or introduce deformable convolutions, and perform classification and regression at different P layers through shared or separate detection heads; First, the global semantic vector or multi-scale feature extracted from the large model is "remapped" to align with the number of channels or spatial resolutions of different scales of FPN. The adaptive attention mechanism is used to weight each scale feature of FPN through the learned weights. An MLP or lightweight network is designed to input the global features output by the large model. The MLP outputs a set of weight vectors (α1, α2, α3, α4, α5) corresponding to the five scales of FPN. When the features are fused, the output of the corresponding layer can be expressed as: P′i=αi×Pi or P′i=αi×Pi+βi×GlobalFeat; During training or verification, the Recall, Precision or mAP of small object detection at each FPN scale is counted. If it is found that the detection effect of a certain scale on small objects is obviously insufficient, it is necessary to increase the weight of this scale in the fusion. According to the monitoring results, αi is corrected twice: αi←αi+Δαi.
8. A process method for detecting tiny objects according to claim 1, characterized in that: Through reinforcement learning agents, adaptive data enhancement parameter tuning is performed and online adaptive updates are achieved, including: Set the model's current small object detection performance indicators, current data enhancement parameters, adjust the data enhancement hyperparameter combination, and give positive or negative rewards; Embed the RL agent into the training loop and call it periodically or at a certain number of steps. The RL agent calculates the reward based on the current detection result, updates the strategy, and outputs the data enhancement parameters for the next stage. If you not only care about small objects, but also large objects or the overall mAP, you can introduce a weighted reward function. The RL agent may do more random exploration in the early stages of training; as the number of epochs increases, the agent will form a stable preference for effective reinforcement strategies, and pass the RL decision results to the data loading and reinforcement pipeline in real time. The performance of the training model in turn updates the RL state, achieving closed-loop feedback.
9. A process method for detecting tiny objects according to claim 1, characterized in that: Perform non-maximum suppression on the detection results, calibrate the confidence, and perform iterative feedback and online fine-tuning, including: Before performing NMS on the detection results, group them according to the size of the prediction box. Set a relatively smaller IoU threshold for small object groups, or use Soft-NMS, DIoU-NMS and other methods to better retain adjacent small targets. For large object groups, keep the regular NMS threshold. If there are dense small objects in the detection scene, use weighted box fusion and other methods instead of traditional hard suppression, and perform intelligent fusion based on the confidence and position deviation of the detection head output; The feature ROI of the candidate box or the corresponding FPN feature is fused again with the global semantic vector extracted by the large model. The confidence is corrected twice through a lightweight MLP or weighting function. For small objects that lack local features but are strongly related to the global context, their confidence is moderately improved. According to the accuracy performance of small and large objects on the validation set, the classification score threshold is dynamically changed, the threshold for small objects is appropriately lowered, and the threshold for large objects is maintained or increased. The samples of small objects that were missed or incorrectly detected in the detection phase are annotated or automatically screened, and the data set is updated online / semi-online to enrich the distribution of small object samples. During the deployment phase, regular fine-tuning is performed based on real-time detection results. During fine-tuning, most of the weights of the backbone and large model are maintained, and the detection head and weights related to small objects are mainly updated to quickly iterate and optimize the small object detection capability.
10. A process method for detecting tiny objects according to claim 1, characterized in that: Deploy large models to the cloud, get real-time global feedback, and continuously update online, including: The cloud server can continuously collect data from different application scenarios and different geographical regions, and perform periodic retraining or incremental training on large models. The local system only needs to regularly upload necessary feature statistics or desensitized detection samples and performance indicators to the cloud. The cloud returns updated pre-trained weights or feature fusion strategies based on global information. The local detection system interacts with the cloud after a set period or when certain trigger conditions are met, and pulls the latest global feature extraction layer weights, large model parameters, and data enhancement strategy recommendations. According to the update package sent by the cloud, the local system fine-tunes or replaces some modules in real time, forming a "cloud-end" collaborative online evolution mechanism; When the local detection scenario changes, the cloud can quickly adapt and generate new parameters after receiving new data.