Negative sample sampling method, negative sample sampling system, device, and medium

CN122574560BActive Publication Date: 2026-09-18STORAGEX TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610992717.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-18
Estimated Expiration
2046-07-06

AI Technical Summary

Technical Problem

[0013]本发明申请提供了负样本采样方法、负样本采样系统、装置、介质,旨在至少部分解决现有目标检测训练过程中负样本数量巨大、空间分布不均、困难负样本学习不足、多尺度特征层负样本监督失衡、类别易混淆背景未被充分利用、采样策略难以随训练阶段自适应调整,以及区域采样后仍可能保留低价值简单负样本等技术问题

Benefits of technology

(1)通过在正样本集合确定后对多尺度特征图中的候选负样本进行局部区域划分,并在局部区域内确定区域候选负样本,可以使各局部区域均具有参与负样本监督的机会,避免全局困难负样本集中在少数局部区域,提高负样本空间分布的均衡性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574560B_ABST
    Figure CN122574560B_ABST
Patent Text Reader

Abstract

The application discloses a negative sample sampling method, a negative sample sampling system, a device and a medium, and belongs to the technical field of computer vision and target detection model training. The negative sample sampling method comprises the following steps: obtaining a candidate negative sample set; dividing candidate prediction positions on each feature map scale into a plurality of local regions, and determining regional candidate negative samples in each local region; constructing a negative sample sampling probability based on a target confidence prediction value corresponding to the regional candidate negative samples and determining a negative sample sampling number; performing negative sample sampling in each local region to obtain a regional sampled negative sample set; performing quality filtering on the regional sampled negative sample set, and performing difficult negative sample backfilling from the candidate negative samples that have not been sampled according to a difficulty score; and constructing a target confidence loss based on positive samples, reserved negative samples and backfilled negative samples. The application can improve the target detection model training efficiency and the false detection suppression capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning model training and target detection model optimization technology, specifically involving negative sample sampling methods, negative sample sampling systems, devices, and media. Background Technology

[0002] Object detection is a crucial task in computer vision, aiming to identify object categories and determine their locations from input images. With the development of deep learning technology, object detection models based on convolutional neural networks, feature pyramid networks, and multi-scale prediction heads have been widely applied in scenarios such as industrial product appearance defect detection, electronic component inspection, surface flaw detection, micro-object detection, and automated quality inspection.

[0003] Existing object detection models typically output prediction results on feature maps of one or more scales after the input image passes through a backbone feature extraction network and a feature fusion network. These prediction results generally include bounding box locations, class predictions, and object confidence scores. The bounding box represents the object's location, the class prediction determines the object's class, and the object confidence score determines whether a candidate location contains the object. During model training, positive and negative samples are determined based on the matching relationship between the ground truth bounding boxes and the candidate predicted locations. Positive samples are typically candidate locations that match the ground truth objects and are used to calculate the bounding box regression loss, classification loss, and positive sample loss for object confidence. Negative samples are typically background candidate locations that do not predict any ground truth objects and are used to calculate negative sample loss for object confidence.

[0004] In one-stage or multi-scale object detection models, the number of candidate predicted locations is typically much larger than the number of true targets. For example, in detection models that use multi-scale feature maps for prediction, a single input image can generate thousands to tens of thousands of candidate predicted locations, while the number of true targets is usually only a few to dozens. Therefore, a severe imbalance between positive and negative samples can easily occur during training. If all background candidate locations are directly used as negative samples in the target confidence loss calculation, a large number of simple background samples will dominate the training gradient, causing the model training process to be more influenced by low-value background samples, thereby reducing training efficiency and weakening the model's ability to learn from high-value, difficult background samples.

[0005] To alleviate the aforementioned problems, various sample allocation or hard sample learning methods have been proposed in the prior art. For example, threshold-based sample allocation methods determine positive and negative samples based on the intersection-union ratio (IU) threshold between predicted and ground truth boxes; dynamic positive sample allocation methods determine the set of positive samples based on the quality of predicted boxes, classification confidence, or dynamic matching strategies; online hard sample mining methods select samples with larger losses or higher confidence from candidate negative samples for training; Focal Loss-like methods can reduce the loss weight of simple samples and increase the relative contribution of hard samples through modulation factors; and random negative sample sampling methods can randomly select a portion of negative samples from all negative samples to reduce the number of negative samples and computational overhead. However, the above-mentioned existing methods still have certain shortcomings, specifically: (1) Existing global hard negative sample mining methods usually select samples with large loss or high confidence from all candidate negative samples. This method tends to concentrate the selected hard negative samples in a few local areas, such as complex texture areas, reflective areas, areas near the target edge or near the annotation boundary, resulting in other background areas lacking effective negative sample supervision, making the model's ability to suppress false detections of the entire image background uneven.

[0006] (2) In the training method that directly uses all background locations to calculate the target confidence loss, a large number of low-confidence, easily distinguishable simple background samples will provide repetitive and low-value gradient information, which not only increases the redundancy of training calculation, but may also reduce the influence of difficult negative samples on the model parameter update, making it difficult for the model to fully learn the discrimination features of high-confidence background, strong texture background and easily misdetectable background.

[0007] (3) Existing global top-k hard negative sample sampling methods may disrupt the supervision balance between multi-scale feature layers. Target detection models typically use feature maps of different scales to detect targets of different sizes. If negative samples are selected based solely on global difficulty, the selected negative samples may be overly concentrated in a certain feature layer, resulting in insufficient supervision of negative samples in small, medium, or large target layers, thus affecting the overall training effect of the multi-scale target detection model.

[0008] (4) Existing negative sample sampling methods often do not make full use of local context information. Background false detections often have spatial correlations; for example, a certain local texture region as a whole is easily misclassified as a target. If negative sample selection is based solely on the prediction confidence of a single candidate location, it is difficult to ensure that each local background region receives adequate supervision, and it is also difficult to balance spatial coverage and learning performance on difficult samples.

[0009] (5) In applications such as industrial defect detection, some background textures, although not real targets, may have high similarity to certain defect categories in terms of category branches, which can easily lead to false positives in the inference stage. Existing methods that select difficult negative samples based solely on target confidence or target confidence loss usually do not make full use of category branch prediction information, and may therefore ignore background candidate positions that are close to the target in terms of category.

[0010] (6) Existing negative sample sampling strategies are usually relatively fixed and difficult to adjust dynamically with the training stage. In the early stage of training, the model's prediction results are still unstable. If too much focus is placed on high-confidence negative samples, noise may be introduced and the training stability may be affected. In the later stage of training, the model has a strong discriminative ability, and it is more necessary to focus on learning high-confidence false positive samples and class confusion backgrounds. Fixed sampling strategies are difficult to balance the stability in the early stage of training and the refinement effect in the later stage of training.

[0011] (7) Existing dynamic positive sample allocation algorithms are mainly used to solve the problem of which candidate positions should be positive samples, but do not specifically solve the problems of negative sample spatial balance, category awareness, difficulty awareness, and adaptive sampling during the training phase. Therefore, even after the positive samples are determined by using dynamic positive sample allocation algorithms, it is still necessary to further design a reasonable negative sample sampling mechanism and a target confidence loss construction method to improve the effectiveness of negative sample supervision.

[0012] (8) Regional equalization sampling alone may still be insufficient. Due to the randomness of regional sampling, some selected negative samples may be simple background samples with low target confidence and small negative sample loss. Although such samples can ensure spatial coverage to a certain extent, their training value is low. If they are directly retained, it will waste the limited negative sample budget; if we completely switch to global hard sample selection, it may destroy the regional equalization effect. Summary of the Invention

[0013] This invention provides a negative sample sampling method, system, apparatus, and medium, aiming to at least partially solve the technical problems in existing target detection training processes, such as the huge number of negative samples, uneven spatial distribution, insufficient learning of difficult negative samples, imbalance in supervision of negative samples in multi-scale feature layers, underutilization of easily confused backgrounds, difficulty in adaptively adjusting sampling strategies during training, and the potential retention of low-value simple negative samples after regional sampling. This invention obtains a candidate negative sample set after determining the positive sample set, performs local region division on the candidate negative samples on the multi-scale feature map, constructs the sampling probability within the region based on the target confidence prediction value corresponding to the candidate negative samples in the region, determines the region sampling negative sample set based on the sampling probability within the region and the number of negative samples sampled, performs quality filtering on the sampled negative samples, and then selects difficult negative samples from the unsampled candidate negative samples for backfilling according to the target confidence loss or difficulty score. Finally, the target confidence loss is constructed based on the positive samples, retained negative samples, and backfilled negative samples. To achieve the objectives of this invention, the following technical solution is adopted: Firstly, a negative sample sampling method is applied to the training of an object detection model, including: Obtain the candidate negative sample set; The candidate prediction locations at each feature map scale are divided into multiple local regions, and regional candidate negative samples belonging to the candidate negative sample set are determined within each local region. Based on the target confidence prediction value corresponding to the candidate negative samples in the region, the negative sample sampling probability in each local region is constructed, and the negative sample sampling quantity corresponding to each local region is determined; according to the negative sample sampling probability and the negative sample sampling quantity, the candidate negative samples in the region are sampled in each local region to obtain the region sampled negative sample set; The negative samples in the negative sample set of the region are quality filtered to remove the sampled negative samples that meet the low value sample condition, and the number of removed negative samples is counted; from the unsampled candidate negative samples, difficult negative samples corresponding to the number of removed negative samples are selected according to the difficulty score and backfilled to obtain the backfilled negative sample set. Based on the positive sample set, the negative sample set of the region sampled after quality filtering, and the backfilled negative sample set, a target confidence loss is constructed for training the target detection model.

[0014] Secondly, a negative sample sampling system, applied to target detection model training, includes: The module retrieves a set of candidate negative samples. The region segmentation module is configured to divide the candidate prediction location at each feature map scale into multiple local regions, and determine the regional candidate negative samples belonging to the candidate negative sample set within each local region. The sampling parameter determination module is configured to construct the sampling probability of negative samples in each local region based on the target confidence prediction value corresponding to the candidate negative samples in the region, and determine the sampling quantity of negative samples corresponding to each local region. The region sampling module is configured to sample candidate negative samples in each local region according to the negative sample sampling probability and the negative sample sampling quantity, thereby obtaining a set of region-sampled negative samples; The quality filtering module is configured to perform quality filtering on negative samples in the negative sample set of the region, remove sampled negative samples that meet the low-value sample condition, and count the number of removed negative samples. The difficult backfilling module is configured to select difficult negative samples from the unsampled candidate negative samples according to the difficulty score, which corresponds to the number of removed negative samples, and backfill them to obtain a backfilled negative sample set. The loss construction module is configured to construct a target confidence loss for training the target detection model based on the positive sample set, the region sampling negative sample set retained after quality filtering, and the backfilled negative sample set.

[0015] Thirdly, an image annotation apparatus includes a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, it executes an image annotation method as described in any of the first aspects above.

[0016] Fourthly, a computer-readable storage medium storing instructions that, when executed on a computer, perform any of the image annotation methods described in the first aspect above.

[0017] Compared with the prior art, the beneficial technical effects of this invention application are as follows: (1) By dividing the candidate negative samples in the multi-scale feature map into local regions after the positive sample set is determined, and determining the regional candidate negative samples in the local regions, each local region can have the opportunity to participate in the supervision of negative samples, avoiding the concentration of global difficult negative samples in a few local regions, and improving the balance of the spatial distribution of negative samples.

[0018] (2) By constructing the negative sample sampling probability within the region based on the target confidence prediction value, the category prediction value and the training phase parameters, it is possible to make the high target confidence background, the category confusion background and the high response background in the later stage of training more likely to be sampled as difficult negative samples, thereby improving the target detection model's ability to learn from false detection backgrounds.

[0019] (3) By determining the number of negative samples corresponding to a local region based on the regional difficulty information, feature map scale parameters and training phase parameters, it is possible to obtain relatively reasonable negative sample supervision for different feature map scales and different local regions, thereby improving the imbalance problem of negative sample supervision in multi-scale feature layers.

[0020] (4) By performing quality filtering on the negative sample set of the region sampling and filling the difficult negative samples back into the unsampled candidate negative samples according to the difficulty score, we can retain the regional balanced sampling effect while avoiding the limited negative sample budget being occupied by low-value simple background samples, so that the negative samples that finally participate in the target confidence loss calculation have both spatial coverage and high training value.

[0021] (5) By constructing the target confidence loss based only on the positive sample set, the negative sample set of the region retained after quality filtering, and the backfilled negative sample set, the redundant loss calculation caused by a large number of simple background samples can be reduced, and the training process can focus more on effective negative samples, thereby improving the training efficiency and false detection suppression ability of the target detection model. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the structure of a negative sample sampling method according to an embodiment of this invention. Figure 2 This is a schematic diagram of the composition of a negative sample sampling system according to an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the composition of an acquisition module according to an embodiment of this invention. The accompanying drawings are provided to further understand the present invention and form part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation thereof. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0025] As shown in the present invention application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate that explicitly identified steps and elements are included, and "multiple" includes one, two, or more, and these steps and elements do not constitute an exclusive list; the method or apparatus may also include other steps or elements.

[0026] While this application makes various references to certain modules of the system according to embodiments of the present invention, any number of different modules can be used and run on user terminals and / or servers. The modules are merely illustrative, and different aspects of the system and method may use different modules.

[0027] This invention application uses flowcharts to illustrate the operations performed by the system according to embodiments of the invention. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously, as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0028] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein.

[0029] It is worth noting that in this invention application, all data acquisition actions are carried out in compliance with the relevant data protection laws and policies of the country where the location is located, and with the authorization granted by the owner of the relevant device.

[0030] Firstly, such as Figure 1 As shown, a negative sample sampling method is applied to the training of an object detection model, including: Obtain the candidate negative sample set; The candidate prediction locations at each feature map scale are divided into multiple local regions, and regional candidate negative samples belonging to the candidate negative sample set are determined within each local region. Based on the target confidence prediction value corresponding to the candidate negative samples in the region, the negative sample sampling probability in each local region is constructed, and the negative sample sampling quantity corresponding to each local region is determined; according to the negative sample sampling probability and the negative sample sampling quantity, the candidate negative samples in the region are sampled in each local region to obtain the region sampled negative sample set; The negative samples in the negative sample set of the region are quality filtered to remove the sampled negative samples that meet the low value sample condition, and the number of removed negative samples is counted; from the unsampled candidate negative samples, difficult negative samples corresponding to the number of removed negative samples are selected according to the difficulty score and backfilled to obtain the backfilled negative sample set. Based on the positive sample set, the negative sample set of regions retained after quality filtering, and the backfilled negative sample set, a target confidence loss is constructed for training the target detection model.

[0031] In this invention application, the negative sample sampling method can be applied to one-stage target detection models, multi-scale target detection models, industrial defect detection models, micro-target detection models, or other detection models that require determining positive and negative samples and calculating target confidence loss during training. For example, the target detection model may include YOLO series detection models, YOLOX series detection models, anchor-free detection models, anchor-based detection models, detection models with feature pyramid structures, or other detection models with multi-scale prediction outputs. The negative sample sampling method in this invention application can be used to perform regionalization, difficulty-aware, and backfilling compensation sampling on candidate negative samples after the positive sample set is determined, in order to reduce the dominance of simple background samples on the training gradient and improve the training value of difficult negative samples.

[0032] In some embodiments, the object detection model can perform feature extraction and multi-scale prediction on the input training image to obtain prediction results at at least one feature map scale. The prediction results can include multiple candidate prediction locations, each corresponding to a target confidence prediction value, a category prediction value, and a bounding box prediction value. The target confidence prediction value can be used to characterize the probability or confidence level that the candidate prediction location contains a target, the category prediction value can be used to characterize the probability or score that the candidate prediction location belongs to each target category, and the bounding box prediction value can be used to characterize the predicted box position corresponding to the candidate prediction location.

[0033] In this invention application, the positive sample set can be determined first based on the real annotation information of the training images and the positive sample allocation algorithm. The positive sample allocation algorithm may include Simplified Optimal Transport Assignment (SimOTA), Optimal Transport Assignment (OTA), Adaptive Training Sample Selection (ATSS), Task-Aligned Assigner (TaskAlignedAssigner), CenterSampling, Intersection over Union Threshold Matching (IoU Threshold Matching), or other algorithms capable of determining the positive sample positions based on the real annotation information and prediction results. This invention application does not limit the specific method of determining the positive sample set, as long as the positive sample positions used to undertake the real target prediction task can be determined from the candidate prediction positions.

[0034] In some embodiments, after determining the positive sample set, positive sample positions can be excluded from the candidate prediction positions of the target detection model based on the positive sample set to obtain a candidate negative sample set. The candidate negative sample set can be understood at least as a set of background candidate positions that are not responsible for predicting the real target. Optionally, when obtaining the candidate negative sample set, at least one of the following can also be excluded: invalid padding positions, positions corresponding to ignored labeled regions, and other positions that do not participate in negative sample supervision, in order to avoid invalid or uncertain positions from participating in subsequent negative sample sampling.

[0035] In some embodiments, candidate prediction locations at each feature map scale can be divided into multiple local regions. Different feature map scales can correspond to different feature map heights, feature map widths, and strides. For each feature map scale, the candidate prediction locations at that scale can be gridded according to a preset region size or an adaptive region size to obtain multiple local regions. A local region can be understood at least as a set of local candidate prediction locations on the feature map, and each local region may include one or more candidate prediction locations.

[0036] In some embodiments, after dividing the local regions, regional candidate negative samples belonging to the candidate negative sample set can be determined within each local region. Specifically, for each local region, it can be determined whether the candidate predicted position in that local region belongs to the candidate negative sample set; if a candidate predicted position belongs to the candidate negative sample set, then that candidate predicted position is determined as a regional candidate negative sample within that local region; if a candidate predicted position is a positive sample position, an invalid position, or an ignored position, then it is not considered a regional candidate negative sample. Thus, subsequent sampling can be performed only at valid background candidate positions.

[0037] In some embodiments, the sampling probability of negative samples in each local region can be constructed based on the target confidence prediction value corresponding to the candidate negative samples in the region. A higher target confidence prediction value generally indicates that the background candidate location is more likely to be misclassified as a target by the target detection model, and the corresponding negative sample has higher training value. Therefore, the target confidence prediction value can be activated to obtain a target confidence score, and then the sampling weight of the candidate negative samples in the region can be determined based on the target confidence score. The sampling weights within the same local region are then normalized to obtain the sampling probability of negative samples corresponding to the candidate negative samples in that local region.

[0038] Furthermore, the sampling probability of negative samples can be determined by combining at least one of the following: class prediction value, training phase parameters, feature map scale parameters, and positive sample neighborhood relationships. For example, when the class prediction value of a candidate negative sample in a certain region shows that it is relatively close to a certain target class, the sampling probability of the candidate negative sample in that region can be increased; when training is in the early stage, the sampling probability distribution can be made relatively smooth to improve training stability; when training is in the middle and late stages, the sampling probability can be made more biased towards difficult negative samples with high target confidence to strengthen the model's ability to suppress false detection background.

[0039] In some embodiments, the number of negative samples for each local region can be determined. The number of negative samples can be a fixed number, or it can be determined based on at least one of the following: the number of candidate negative samples within the local region, the mean target confidence score, the maximum target confidence score, the number of high-confidence candidate negative samples, the feature map scale, and training phase parameters. For local regions with high target confidence, many candidate negative samples, or a higher likelihood of false positives, a larger number of negative samples can be configured; for local regions with low target confidence or few candidate negative samples, a smaller number of negative samples can be configured. This balances the spatial coverage and training value of the negative samples.

[0040] In some embodiments, candidate negative samples can be sampled in each local region according to the negative sample sampling probability and the number of negative samples, resulting in a set of region-sampled negative samples. Specifically, for each local region, sampling without replacement can be performed according to the negative sample sampling probability corresponding to the candidate negative samples in that local region, ensuring that the number of samples does not exceed the number of candidate negative samples in that local region. The negative samples obtained from each local region are then merged to obtain the set of region-sampled negative samples. By sampling separately in each local region, it is possible to avoid negative samples being concentrated in only a few high-response areas, ensuring that different spatial regions have a chance to provide negative sample supervision.

[0041] In some embodiments, after obtaining the set of negative samples from the region sampling, quality filtering can be performed on the negative samples in the set. Quality filtering can be used to identify and remove sampled negative samples with low training value. Specifically, a quality score can be calculated for each negative sample in the set of negative samples from the region sampling. The quality score can be determined based on at least one of the following: target confidence score, target confidence loss, or other metrics that can characterize the training value of the negative sample. When the quality score of a sampled negative sample is lower than a preset quality threshold, the sampled negative sample can be considered a low-value sample and removed from the set of negative samples from the region sampling.

[0042] In some embodiments, the low-value sample condition may include at least one of the following: a target confidence score lower than a preset confidence threshold, a target confidence loss lower than a preset loss threshold, and a difficulty score lower than a preset difficulty threshold. The preset quality threshold can be a fixed threshold or can be adjusted based on at least one of the following: training stage, feature map scale, local region difficulty, and historical false positive statistics. Quality filtering can minimize the selection of a large number of simple background samples due to randomness during region sampling, thus preventing low-value samples from consuming the limited negative sample budget.

[0043] In some embodiments, the difficulty score can be determined based on at least one of the following: target confidence loss, target confidence prediction, target confidence score, class confusion score, and class entropy. For example, the target confidence loss when the target is the background can be used as the difficulty score. A higher target confidence loss corresponding to an unsampled candidate negative sample indicates that the background candidate position is more likely to be misclassified as the target by the model, resulting in a higher difficulty score and making it suitable as a backfill negative sample. Thus, low-value negative samples that are filtered out can be compensated by high-value difficult negative samples from the unsampled candidate negative samples.

[0044] In some embodiments, after removing sampled negative samples that meet the low-value sample criteria, the number of removed negative samples can be counted. The number of removed negative samples can be used to determine the number of difficult negative samples that need to be backfilled subsequently. That is, when several negative samples in the regional sampled negative sample set are filtered out due to low value, the corresponding number can be replenished by using difficult negative sample backfilling to keep the size of the negative samples participating in the target confidence loss calculation relatively stable.

[0045] In some embodiments, difficult negative samples corresponding to the number of removed negative samples can be selected from the unsampled candidate negative samples for backfilling, based on their difficulty scores, to obtain a backfilled negative sample set. Unsampled candidate negative samples can be understood as candidate negative samples that belong to the candidate negative sample set but have not entered the regional sampling negative sample set. For these candidate negative samples, their difficulty scores can be calculated and sorted from high to low. Candidate negative samples with the highest rankings and a quantity corresponding to the number of removed negative samples are selected as backfilled negative samples.

[0046] In some embodiments, if the number of unsampled candidate negative samples is less than the number of removed negative samples, all unsampled candidate negative samples can be identified as backfill negative samples. Furthermore, to avoid backfill negative samples being overly concentrated at a certain feature map scale or a certain local region, an upper limit can be set for the backfill ratio, the feature map scale backfill limit, or the local region backfill limit, so as to maintain the spatial balance of negative sample supervision while supplementing difficult negative samples.

[0047] In some embodiments, the target confidence loss for training the object detection model can be constructed based on a set of positive samples, a set of negative samples from regions retained after quality filtering, and a set of backfilled negative samples. Specifically, candidate prediction positions in the set of positive samples can be used as positive samples for target confidence, and candidate prediction positions in the set of negative samples from regions retained after quality filtering and the set of backfilled negative samples can be used as negative samples for target confidence. The target confidence loss is then calculated only based on these positive and negative samples. This avoids all background candidate positions from participating in the calculation of the target confidence loss, reducing the impact of a large number of simple background samples on gradient updates.

[0048] In some embodiments, the target confidence loss can employ binary cross-entropy loss, weighted binary cross-entropy loss, Focal Loss, or other loss functions suitable for target confidence supervision. For example, for candidate predicted positions in the positive sample set, the target confidence target value can be set to 1; for candidate predicted positions in the region-sampled negative sample set retained after quality filtering and the backfilled negative sample set, the target confidence target value can be set to 0. By constructing the target confidence loss based on the positive sample set, the region-sampled negative sample set retained after quality filtering, and the backfilled negative sample set, the target detection model can focus on learning effective positive samples and high-value negative samples during training.

[0049] Therefore, in the negative sample sampling method of this invention, after obtaining the candidate negative sample set, the candidate prediction positions on the multi-scale feature map are divided into multiple local regions, and negative sample sampling is performed in each local region based on the target confidence prediction value. This makes the negative sample supervision more spatially balanced and avoids difficult negative samples from being concentrated in a few local regions. By performing quality filtering on the region-sampled negative sample set, simple background negative samples with low training value can be removed, reducing the occupation of the negative sample budget by low-value samples. By backfilling the unsampled candidate negative samples according to the difficulty score, more valuable difficult negative samples can be added. By constructing the target confidence loss based on the positive sample set, the region-sampled negative sample set retained after quality filtering, and the backfilled negative sample set, the false detection suppression ability of the target detection model for difficult backgrounds, strong texture backgrounds, easily confused category backgrounds, and target edge backgrounds can be improved, and the negative sample utilization efficiency during the target detection model training process can be improved.

[0050] Optionally, the candidate prediction positions at each feature map scale are divided into multiple local regions, including: obtaining the feature map height, feature map width, and region size corresponding to each feature map scale; dividing the candidate prediction positions at the corresponding feature map scale into multiple local regions according to the region size; when the feature map height or feature map width cannot be divided by the region size, padding is performed on the candidate prediction positions at the feature map scale, and the padded positions are marked as invalid positions; when determining the regional candidate negative samples within each local region, positive sample positions and invalid positions are excluded; the region size is a preset fixed size or the region size should be determined according to at least one of the step size of the corresponding feature map scale, feature map size, input image size, target scale statistics, and training phase parameters.

[0051] In some embodiments, before performing region sampling on candidate negative samples, the feature map height, feature map width, and region size corresponding to each feature map scale output by the object detection model can be obtained first. The feature map height and width can be used to determine the spatial distribution range of candidate predicted positions at the current feature map scale, and the region size can be used to determine the granularity of local region partitioning at that feature map scale. Since object detection models can typically output candidate predicted positions at multiple feature map scales, regions can be partitioned separately for each feature map scale, so that candidate negative samples at different scales can all receive regionalized supervision within their corresponding feature map space.

[0052] In some embodiments, each feature map scale can correspond to a different step size. For example, the object detection model can output prediction results at three feature map scales with step sizes of 8, 16, and 32. For feature map scales with smaller step sizes, the feature map resolution is higher, and the number of candidate prediction locations is typically larger; for feature map scales with larger step sizes, the feature map resolution is lower, and the number of candidate prediction locations is typically smaller. Therefore, when performing local region segmentation, the corresponding region size can be determined for different feature map scales to ensure that the region segmentation results can adapt to the distribution and spatial density of candidate prediction locations at different scales.

[0053] In some embodiments, candidate prediction locations at the corresponding feature map scale can be divided into multiple local regions according to region size. Specifically, based on the feature map height and width at the current feature map scale, the region can be divided into blocks in both the height and width directions according to the region size, thereby obtaining multiple local regions. Each local region may include one or more candidate prediction locations within that region. In this way, candidate prediction locations that are originally distributed across the entire feature map scale can be divided into multiple local regions, allowing subsequent negative sample sampling to be performed separately on a local region basis.

[0054] In some embodiments, if the feature map height is H, the feature map width is W, and the region size is R at the current feature map scale, then the number of local regions in the height direction can be determined as H / R (rounded up), and the number of local regions in the width direction can be determined as W / R (rounded up). That is, when the feature map height or feature map width is not divisible by the region size, edge regions can still be retained by rounding up to avoid missing candidate prediction positions at the feature map edges. Therefore, all valid candidate prediction positions at the current feature map scale can be included in the corresponding local regions.

[0055] In some embodiments, when the feature map height or width is not divisible by the region size, candidate prediction positions at that feature map scale can be padded. Specifically, several positions can be padded in the height direction, width direction, or both directions of the feature map, so that the height and width of the padded feature map are divisible by the region size. This facilitates subsequent batch, tensor, or matrix-based statistical analysis of candidate negative samples, probability calculation, and determination of the number of negative samples. The padded positions do not correspond to the actual candidate prediction positions output by the target detection model. Accordingly, the padded positions can be marked as invalid positions. Invalid positions can be recorded using invalid position masks, padding markers, or other state marking methods, thereby preventing padded positions from being mistakenly identified as actual candidate prediction positions and participating in negative sample sampling, thus ensuring the regularity of region division and the validity of sampling results.

[0056] In some embodiments, when determining candidate negative samples within each local region, positive sample positions and invalid positions can be excluded. Specifically, for a candidate prediction position in a certain local region, it can be first determined whether the candidate prediction position belongs to the positive sample set; if the candidate prediction position belongs to the positive sample set, then the candidate prediction position is used to undertake the real target prediction task and is not used as a negative sample sampling object. Further, it can be determined whether the candidate prediction position belongs to an invalid position obtained by completion; if it is an invalid position, then the position does not correspond to the real prediction result and is not used as a negative sample sampling object. Thus, only candidate prediction positions that are neither positive sample positions nor invalid positions, and that belong to the candidate negative sample set, are determined as candidate negative samples within the corresponding local region.

[0057] In other embodiments, in addition to excluding positive sample locations and invalid locations, locations corresponding to ignored regions may also be excluded. Ignored regions may include manually annotated ignored regions, low-quality annotated boundary regions, regions whose belonging to the target cannot be determined, or other locations unsuitable for negative sample supervision. By excluding the above-mentioned locations during the determination of candidate negative samples in the region, the interference of uncertain or invalid samples on the negative sample sampling results can be reduced, making the set of candidate negative samples in the region more suitable as the target confidence negative sample supervision object.

[0058] In some embodiments, the region size can be a preset fixed size, and a mapping relationship between feature map scales and region sizes can be established in advance. For example, the region size at each feature map scale can be set to 2, so that each local region covers 2×2 candidate prediction locations at the corresponding feature map scale; alternatively, different preset fixed region sizes can be set for different feature map scales, such as using a larger region size at high-resolution feature map scales and a smaller region size at low-resolution feature map scales. Thus, the granularity of local region partitioning can be flexibly configured according to the number and spatial distribution differences of candidate prediction locations at different feature map scales.

[0059] In other embodiments, the region size can also be determined based on at least one of the following: step size of the corresponding feature map scale, feature map size, input image size, target scale statistics, and training phase parameters. For example, the region size can be dynamically calculated based on the step size of the feature map scale, feature map height, feature map width, input image size, and training progress, so that the granularity of local region division at different feature map scales matches the spatial distribution of candidate prediction locations.

[0060] For example, when the step size of a certain feature map scale is small, the feature map size is large, and the number of candidate prediction positions is large, the region size can be appropriately increased to allow a single local region to cover more candidate prediction positions, thereby reducing the number of local regions and reducing the computational cost of region sampling. When the step size of a certain feature map scale is large, the feature map size is small, and the number of candidate prediction positions is small, the region size can be appropriately reduced to make the local region division more refined and avoid the single local region covering too large an area, which would weaken the selection ability of local difficult negative samples.

[0061] In some embodiments, the region size can be determined based on target scale statistics. Target scale statistics may include the average size, median size, scale distribution, size distribution of different target categories in the training dataset, target size distribution in the current batch of training images, or the target size range to be detected at each feature map scale. When there are many small targets in the training task, a finer region division granularity can be set at the high-resolution feature map scale to enhance negative sample supervision of the background region around the small target, the target edge region, and the locally high-response background region; when there are many large targets in the training task, a region size adapted to the prediction range of large targets can be set at the low-resolution feature map scale to improve the balance of negative sample supervision of the background region under the large target scale.

[0062] In some embodiments, the region size can also be dynamically adjusted based on training phase parameters. Training phase parameters may include the current training epoch, the current iteration count, the total number of training epochs, the total number of iterations, or the training progress percentage. In the early stages of training, the prediction results of the object detection model are not yet stable, so a relatively smooth region partitioning strategy can be adopted to ensure that each local region has a relatively balanced opportunity to sample negative samples, thereby improving training stability. In the later stages of training, the object detection model has already acquired a certain discriminative ability, and the region size can be appropriately adjusted so that local regions can more effectively cover high-confidence backgrounds, target edge backgrounds, strong texture backgrounds, or backgrounds that are easily confused by categories, thereby improving the learning effect of difficult negative samples.

[0063] Therefore, the region size can be determined using a fixed mapping method, an adaptive calculation method, or a combination of both. The fixed mapping method involves pre-configuring different region sizes for feature map scales of different lengths. The adaptive calculation method calculates the region size based on the feature map size, input image size, target scale statistics, and training progress ratio, and imposes upper and lower limits on the calculated region size to avoid excessively large region sizes leading to coarse local supervision, or excessively small region sizes leading to an excessive number of regions and increased sampling computation. This allows the region size to adapt to both the spatial structure of multi-scale feature maps and the negative sample sampling requirements under different training stages and target scale distributions.

[0064] It is understandable that by obtaining the feature map height, feature map width, and region size corresponding to each feature map scale, and dividing the candidate prediction positions into local regions according to the region size, the negative sample sampling process can be transformed from global candidate negative sample selection to negative sample selection within multi-scale local regions. By padding when the feature map height or feature map width is not divisible by the region size and marking the padded positions as invalid positions, invalid positions can be avoided from participating in sampling while ensuring the regularity of region division. By excluding positive sample positions and invalid positions when determining regional candidate negative samples, it can be ensured that all subsequent sampling objects are valid candidate negative samples.

[0065] Therefore, by dividing the feature map into local regions at each feature map scale, on the one hand, candidate negative samples at different feature map scales can participate in sampling according to the spatial structure of the corresponding scale, avoiding excessive concentration of negative sample supervision at a certain scale or certain local high-response regions; on the other hand, the region size can be determined according to at least one of the step size, feature map size, input image size, target scale statistics, and training stage parameters, so that the granularity of region division can adapt to the needs of different detection models, different target scale distributions, and different training stages, thereby improving the spatial balance, multi-scale adaptability, and training stability of negative sample sampling.

[0066] Optionally, the target confidence prediction value corresponding to the candidate negative sample in the region is activated to obtain a target confidence score; based on the target confidence score and the difficulty adjustment parameter, the basic sampling weight of the candidate negative sample in the region is determined; the basic sampling weight of each candidate negative sample in the same local region is processed to obtain the negative sample sampling probability in the corresponding local region; wherein, the difficulty adjustment parameter changes dynamically with the training progress to make the negative sample sampling probability distribution in the early stage of training smoother and to make the negative sample sampling probability in the later stage of training more biased towards difficult negative samples with high target confidence.

[0067] In some embodiments, after determining the candidate negative samples within each local region, the sampling probability of the negative samples can be further constructed based on the target confidence prediction value corresponding to each candidate negative sample. The target confidence prediction value can be understood at least as the target existence prediction result output by the target detection model for the candidate prediction location. For negative samples, their true target confidence label is usually the background label; if the target confidence prediction value corresponding to a candidate negative sample in a certain region is high, it indicates that the candidate background location is more likely to be misclassified as a target by the target detection model, and the candidate negative sample in that region has higher training value than the background sample with low target confidence.

[0068] In some embodiments, the target confidence prediction value corresponding to the candidate negative sample in the region can be activated first to obtain a target confidence score. For example, when the target confidence prediction value output by the target detection model is a target confidence logistic value, the Sigmoid activation function can be used to activate this target confidence logistic value to obtain a target confidence score ranging from 0 to 1. The closer the target confidence score is to 1, the easier it is for the candidate negative sample in the region to be predicted as a target by the model; the closer the target confidence score is to 0, the easier it is for the candidate negative sample in the region to be identified as background by the model.

[0069] In some embodiments, after obtaining the target confidence score, the basic sampling weights for candidate negative samples in a region can be determined based on the target confidence score and a difficulty adjustment parameter. Specifically, the target confidence score can be used as a basic measure of the difficulty of negative samples, and the sensitivity of the sampling weights to the target confidence score can be adjusted by the difficulty adjustment parameter. When the difficulty adjustment parameter is small, the difference in basic sampling weights corresponding to different target confidence scores is small, making the sampling of negative samples in local regions closer to uniform sampling; when the difficulty adjustment parameter is large, the basic sampling weights corresponding to high target confidence scores will be further amplified, making it easier to select difficult negative samples with high target confidence during the sampling process.

[0070] In some embodiments, the basic sampling weights can be determined using exponentiation. For example, the target confidence score can first be truncated to a lower limit to avoid the sampling weights being zero or unstable due to excessively low target confidence scores. Then, the truncated target confidence score is exponentially calculated according to the difficulty adjustment parameter to obtain the basic sampling weights. That is, the higher the target confidence score, the higher the basic sampling weights; as the difficulty adjustment parameter increases, the difference in sampling weights between candidate negative samples in high-target-confidence regions and candidate negative samples in low-target-confidence regions further increases.

[0071] In some embodiments, the difficulty adjustment parameter can be dynamically changed with the training progress. The training progress can be determined based on the current training round, the current number of iterations, the total number of training rounds, the total number of iterations, or the proportion of training progress. For example, a smaller difficulty adjustment parameter can be set in the early stages of training to make the negative sample sampling probability distribution relatively smooth, avoiding over-focusing on a few high-confidence negative samples when the model prediction is not yet stable, thereby reducing the impact of noisy negative samples on the training process; in the middle and later stages of training, the difficulty adjustment parameter can be gradually increased to make the negative sample sampling probability gradually biased towards difficult negative samples with high target confidence, thereby enhancing the model's ability to learn from high-confidence false detection backgrounds.

[0072] In some embodiments, the difficulty adjustment parameter can adopt linear growth, piecewise growth, cosine variation, or other scheduling methods that change with the training progress. For example, at the beginning of training, the difficulty adjustment parameter can be at a preset minimum value, ensuring that all candidate negative samples within a local area have a certain sampling opportunity; as training progresses, the difficulty adjustment parameter can be gradually increased; in the later stages of training, the difficulty adjustment parameter can approach a preset maximum value, making the sampling process focus more on background candidate locations with higher target confidence scores and a greater risk of false detection. Thus, the negative sample sampling strategy can be stable in the early stages of training and have the ability to focus on difficult samples in the later stages.

[0073] In some embodiments, after determining the basic sampling weights of candidate negative samples in each region within the same local area, the basic sampling weights within the same local area can be processed to obtain the negative sample sampling probability of the corresponding local area. Specifically, the basic sampling weights of all candidate negative samples in the same local area can be summed, and the basic sampling weight of each candidate negative sample can be divided by the sum of these weights to obtain the negative sample sampling probability of that candidate negative sample in that local area. Thus, the sum of the negative sample sampling probabilities corresponding to candidate negative samples in each region within the same local area can be 1, which facilitates subsequent sampling within the region based on the sampling probability.

[0074] In some embodiments, if the sum of the base sampling weights of all candidate negative samples within a certain local region is zero, or the sum of the base sampling weights is less than a preset stability threshold, then a uniform sampling probability can be applied to the candidate negative samples within that local region, or a preset smoothing term can be added to the base sampling weights before normalization. By setting a smoothing term or reverting to a uniform sampling probability, situations where the sampling probability cannot be stably calculated due to a low target confidence score, numerical underflow, or a small number of candidate negative samples within a local region can be avoided, thereby improving the numerical stability of the negative sample sampling process.

[0075] In some embodiments, the negative sample sampling probability is obtained by independent normalization within each local region. That is, different local regions do not directly compete for the same global sampling probability; instead, they determine their relative sampling probabilities within their respective local regions based on the target confidence score and difficulty adjustment parameters. This approach avoids high-confidence background regions occupying too many sampling opportunities, allowing each local region to sample based on the difficulty of its own candidate negative samples, thereby improving the spatial balance of negative sample sampling.

[0076] In some embodiments, the construction of the negative sample sampling probability can be used in conjunction with the negative sample sampling quantity. For a certain local region, the negative sample sampling probability can be determined first based on the target confidence prediction value and difficulty adjustment parameter of the candidate negative samples in each region within the region. Then, negative samples are selected from the local region based on the corresponding negative sample sampling quantity. Thus, the negative sample sampling probability is used to determine which candidate negative samples are more likely to be selected within the same local region, and the negative sample sampling quantity is used to determine how many negative samples the local region contributes to the final region's negative sample set. Together, they achieve the selection of difficult negative samples within the local region.

[0077] This invention constructs a target confidence score, basic sampling weights, local normalization processing, and a dynamic difficulty adjustment mechanism during the training phase. This avoids the imbalance in negative sample distribution caused by all negative samples or globally high-confidence negative samples directly participating in training. By activating the target confidence prediction value, the model output can be converted into a target confidence score that can characterize the risk of false detection in the background. By determining the basic sampling weights based on the target confidence score and the difficulty adjustment parameter, high-target-confidence background samples have a higher probability of being selected during the sampling process, enhancing the target detection model's ability to suppress high-confidence backgrounds, complex texture backgrounds, and regions prone to false detection. By processing the basic sampling weights within the same local region, a negative sample sampling probability applicable to sampling within the region can be obtained. By dynamically changing the difficulty adjustment parameter with the training progress, sampling stability can be considered in the early stages of training, while the sampling intensity of difficult negative samples can be increased in the later stages of training, improving the adaptability of the negative sample sampling strategy to different training stages.

[0078] Optionally, the basic sampling weights of each candidate negative sample within the same local region are processed to obtain the negative sample sampling probability within the corresponding local region, including: Based on the predicted class value corresponding to the candidate negative sample in the region, a class-aware weight is determined, which is determined according to at least one of the maximum class probability, class entropy, class confusion score, and class prior weight; based on the neighborhood relationship between the candidate negative sample and the positive sample, a positive sample neighborhood modulation weight is determined; based on the feature map scale corresponding to the candidate negative sample in the region, a scale modulation weight is determined. Based on the basic sampling weight, and at least one of the category-aware weight, the positive sample neighborhood modulation weight, and the scale modulation weight, the comprehensive sampling weight of the candidate negative sample in the region is obtained, and the sampling probability of the negative sample is obtained based on the comprehensive sampling weight.

[0079] In some embodiments, category-aware weights can be determined by combining the category prediction values ​​corresponding to candidate negative samples in a region. The category prediction value can be understood at least as the category branch prediction result output by the object detection model for the candidate prediction location. For negative samples, their true semantics are usually background, but some background regions may exhibit prediction features similar to certain target categories in the category branch. For example, scratch textures, reflective areas, dirty areas, edge transition areas, or regular texture areas in industrial defect detection scenarios may be close to certain defect categories in category prediction. Although such backgrounds do not contain real targets, they are more likely to generate false positives during the inference stage; therefore, category-aware weights can be used to increase the probability of their sampling.

[0080] In some embodiments, the category-aware weight can be determined based on the maximum category probability. Specifically, the predicted category values ​​corresponding to candidate negative samples in a region can be activated to obtain the category probabilities for each category, and the maximum category probability can be selected from these probabilities. The higher the maximum category probability, the more likely the candidate negative sample in that region is to be classified as a target category by the category branch, thus indicating that the background location has a higher risk of category confusion. Therefore, the category-aware weight of the corresponding candidate negative sample in the region can be increased based on the maximum category probability, making background candidate locations that are more similar to the target in terms of category easier to sample.

[0081] In other embodiments, the category-aware weights can also be determined based on category entropy. Category entropy can be used to characterize the uncertainty of a region's candidate negative sample in terms of category branching. When the category entropy is high, it indicates that the model is less certain about the category judgment of the candidate position, and the candidate negative sample may be located at a category boundary, a complex texture region, or an easily confused background region. When the category entropy is low and the maximum category probability is high, it indicates that the model may more clearly misclassify the background candidate position as a certain target category. Therefore, category entropy can be used as a component of the category-aware weights, giving background negative samples with high category uncertainty or high risk of category confusion a higher sampling opportunity.

[0082] In some other embodiments, the category-aware weights can also be determined based on the category confusion score. The category confusion score can be determined based on the category false detection rate statistically obtained during the current training process, the confusion matrix of different categories during historical validation, the prior configuration of easily confused categories in business contexts, or the category probability distribution corresponding to regional candidate negative samples. For example, when a certain category has a high false detection rate during historical training or validation, a higher category confusion score can be assigned to the background candidate positions associated with that category, enabling the model to learn that easily confused background more fully in subsequent training.

[0083] In other embodiments, the category-aware weights can also be determined based on category prior weights. Category prior weights can be pre-set based on the false detection cost, category importance, number of category samples, or category detection difficulty for different categories in the business scenario. For example, in an industrial defect detection scenario, higher category prior weights can be set for defect categories with high false detection costs or those more easily confused with background textures; when a candidate negative sample in a region is close to a high false detection cost category in the category branch, its category-aware weight can be increased. Thus, the negative sample sampling strategy can be adapted to the false detection risk in a specific business scenario.

[0084] In some embodiments, the neighborhood modulation weights of positive samples can be determined based on the neighborhood relationships between candidate negative samples and positive sample locations. A positive sample location can be understood at least as a candidate prediction location determined by the set of positive samples. The neighborhood relationships between candidate negative samples and positive sample locations can include spatial distance relationships, center point distance relationships, boundary proximity relationships, relationships within the same local region, relationships between adjacent local regions, or adjacency relationships at the same feature map scale. Since regions near target edges, surrounding background areas, or labeled boundaries are more prone to misjudging target confidence during training and inference, their sampling weights can be adjusted based on the proximity between candidate negative samples and positive sample locations.

[0085] In other embodiments, the modulation weight of the positive sample neighborhood can be determined based on the distance from the candidate negative sample in the region to the nearest positive sample location. For example, the grid distance, Euclidean distance, Manhattan distance, or other spatial distance between the candidate negative sample in the region and the nearest positive sample location can be calculated on the same feature map scale. When the candidate negative sample in the region is close to the positive sample location, it can be considered to be located near the target neighborhood or target boundary, and the modulation weight of the positive sample neighborhood is increased accordingly; when the candidate negative sample in the region is far from the positive sample location, it can be considered to be more likely to belong to a normal background region, and the modulation weight of the positive sample neighborhood is decreased or maintained accordingly.

[0086] In some embodiments, the positive sample neighborhood modulation weights can be used to form a center protection mode or a boundary enhancement mode. In center protection mode, the sampling weights of negative samples that are too close to the center of the positive sample can be reduced to avoid applying overly strong negative sample supervision to candidate locations near the target center. In boundary enhancement mode, the sampling weights of negative samples located near the positive sample boundary or outside the positive sample neighborhood can be increased to enhance the model's ability to distinguish between the target edge background and the confusing background around the target. These modes can be selected and configured according to the characteristics of the detection task, annotation quality, and false detection distribution.

[0087] In some embodiments, scale modulation weights can be determined based on the feature map scale corresponding to the candidate negative samples in a region. Different feature map scales are typically responsible for predicting targets of different sizes; high-resolution feature map scales typically correspond to small target detection, while low-resolution feature map scales typically correspond to large target detection. Since the number of candidate prediction locations, background complexity, negative sample density, and false detection risk may differ at different feature map scales, scale modulation weights can be used to balance the contribution of each feature map scale to the negative sample sampling results.

[0088] In other embodiments, the scale modulation weights can be determined based on at least one of the following: the step size of the feature map scale, the feature map resolution, the number of candidate negative samples, the historical false detection distribution, and the target scale distribution. For example, when a certain feature map scale is responsible for detecting small targets and there are many false background detections at that scale, the scale modulation weight at that feature map scale can be increased to give difficult background negative samples within that scale more sampling opportunities. When the number of candidate negative samples at a certain feature map scale is small or the false detection risk is low, the corresponding scale modulation weight can be reduced or maintained to avoid that scale having a disproportionate impact on the final negative sample set.

[0089] In some embodiments, based on the basic sampling weights determined according to the target confidence prediction value, a comprehensive sampling weight for regional candidate negative samples can be obtained according to the basic sampling weights and at least one of the category-aware weights, positive sample neighborhood modulation weights, and scale modulation weights. For example, the basic sampling weights can be multiplied by the category-aware weights to obtain a comprehensive sampling weight that simultaneously considers the difficulty of achieving target confidence and the risk of category confusion; alternatively, it can be further multiplied by the positive sample neighborhood modulation weights and scale modulation weights so that the comprehensive sampling weight simultaneously reflects the target confidence, the degree of category confusion, the positive sample neighborhood relationship, and the feature map scale difference.

[0090] In some embodiments, the comprehensive sampling weights can also be obtained using weighted summation, weighted product, normalization fusion, or other weight fusion methods. For example, fusion coefficients can be configured for the basic sampling weights, category-aware weights, positive sample neighborhood modulation weights, and scale modulation weights respectively, and these coefficients can be dynamically adjusted according to different training stages. In the early stages of training, the influence of category-aware weights and positive sample neighborhood modulation weights can be appropriately reduced to improve sampling stability; in the later stages of training, the influence of category-aware weights and positive sample neighborhood modulation weights can be increased, making negative sample sampling pay more attention to easily confused backgrounds and target edge backgrounds.

[0091] In some embodiments, after obtaining the comprehensive sampling weights of candidate negative samples in each region within the same local area, the comprehensive sampling weights within that local area can be normalized to obtain the negative sample sampling probability. Specifically, the comprehensive sampling weight of a candidate negative sample in a certain region can be divided by the sum of the comprehensive sampling weights of all candidate negative samples in the same local area to obtain the negative sample sampling probability of that candidate negative sample in the corresponding local area. If the sum of the comprehensive sampling weights is zero or less than a preset stability threshold, a smoothing term can be added before normalization, or the process can be reverted to a uniform sampling probability to improve the stability of the sampling probability calculation process.

[0092] In some embodiments, the negative sample sampling probability can be calculated independently within each local region. That is, the overall sampling weight within each local region is normalized only with the candidate negative samples within that local region, without competing globally with candidate negative samples in other local regions. In this way, while maintaining the spatial balance of local regions, it is easier to select difficult negative samples within each local region that have higher target confidence, are more easily confused by the class, are located near the neighborhood of positive samples, or whose scale requires stronger supervision.

[0093] In this invention application, by introducing category-aware weights, negative samples from the background that are closer to the target category or have higher category uncertainty on the category branch can obtain a higher sampling opportunity; by introducing positive sample neighborhood modulation weights, the supervision of negative samples near the target edge or in the target neighborhood background can be enhanced; by introducing scale modulation weights, the number of candidate negative samples and the difference in false detection risk at different feature map scales can be adapted; by extending the basic sampling weights through category-aware weights, positive sample neighborhood modulation weights, and scale modulation weights, while considering category confusion, spatial neighborhood relationships, and multi-scale feature differences, the probability of sampling easily confused backgrounds, target edge backgrounds, and multi-scale high-risk backgrounds can be increased, enabling the target detection model to learn more fully from negative samples that are prone to false detection during training; and while maintaining the local region sampling framework, category information, spatial information, and scale information can be integrated into the negative sample sampling probability construction process, thereby improving the targeting of negative sample sampling and the target detection model's ability to suppress false detections in complex backgrounds.

[0094] Optionally, determining the number of negative samples corresponding to each local region includes: counting at least one of the following region difficulty information in each local region: the number of regional candidate negative samples, the number of high-confidence candidate negative samples, the mean target confidence score, and the maximum target confidence score; and determining the number of negative samples corresponding to each local region based on at least one of the following region difficulty information, feature map scale parameters, and training phase parameters; wherein the number of negative samples is not greater than the number of regional candidate negative samples in the corresponding local region. And / or, sampling candidate negative samples in each of the local regions according to the negative sample sampling probability and the negative sample sampling quantity includes: sampling candidate negative samples in each of the local regions without replacement according to the negative sample sampling probability.

[0095] In this invention application, the negative sample sampling quantity can be understood at least as the number of negative samples allowed to be selected into the negative sample set for each local region. Since the number of candidate negative samples, the target confidence response level, and the false detection risk are not the same in different local regions, using the same sampling quantity for all local regions may result in simple background regions consuming too much negative sample budget, or difficult background regions not receiving sufficient supervision. Therefore, this invention application can adaptively determine the negative sample sampling quantity based on the regional difficulty information of the local region.

[0096] In some embodiments, the number of candidate negative samples within each local region can be counted. The number of candidate negative samples can be used to represent the size of the effective negative samples that can be sampled within the current local region. When the number of candidate negative samples in a local region is large, it indicates that the local region has a large sampleable space, and a corresponding number of negative samples can be allocated according to the difficulty of the region; when the number of candidate negative samples in a local region is small, it is necessary to limit the number of negative samples to be sampled in that local region to avoid the number of samples exceeding the actual number of negative samples that can be sampled.

[0097] In some embodiments, the number of high-confidence candidate negative samples within each local region can be counted. High-confidence candidate negative samples can be candidate negative samples from regions where the target confidence score is greater than a preset confidence threshold. For object detection models, the true semantics of negative samples are background. If some background candidate locations still have high target confidence, it indicates that these candidate locations are more likely to be misclassified as targets by the model. Therefore, the number of high-confidence candidate negative samples can be used to characterize the number of high-risk false positive backgrounds within a local region. The more high-confidence candidate negative samples there are, the higher the regional difficulty of that local region is generally, and correspondingly, more negative samples can be allocated for sampling.

[0098] In some embodiments, the mean target confidence score can be calculated for each local region. The mean target confidence score can be used to reflect the overall target response level of candidate negative samples within that local region. When the mean target confidence score is high within a certain local region, it indicates that the overall background of that region is more easily identified as a target by the target detection model, which may correspond to a strong textured background, a reflective background, a dense background, a background near the target edge, or a background with easily confused categories. In this case, the number of negative samples for that local region can be appropriately increased to provide more sufficient negative sample supervision for that local region.

[0099] In some embodiments, the maximum target confidence score can be calculated for each local region. The maximum target confidence score can be used to determine whether there are single or a small number of high-response background candidate locations within a local region. Even if the mean target confidence score of a certain local region is low, the presence of candidate negative samples with a high maximum target confidence score within that local region may indicate the presence of background locations prone to false detection. Therefore, a certain number of samples can be reserved for that local region based on the maximum target confidence score to avoid missing locally high-confidence, difficult negative samples.

[0100] Therefore, by statistically analyzing the number of candidate negative samples, the number of high-confidence candidate negative samples, the mean target confidence level, and the maximum target confidence level, the difficulty of negative samples in a local region can be assessed from the perspectives of candidate size, overall response level, and local highest response. Correspondingly, the regional difficulty information can be determined based on at least one of the following: the number of candidate negative samples, the number of high-confidence candidate negative samples, the mean target confidence level, and the maximum target confidence level. For example, the number of high-confidence candidate negative samples, the mean target confidence level, and the maximum target confidence level can be weighted and fused to obtain a regional difficulty score; alternatively, separate judgment rules can be set, such as increasing the sampling quantity when the maximum target confidence level is greater than a preset threshold, or decreasing the sampling quantity when the number of high-confidence candidate negative samples is zero and the mean target confidence level is low. In this way, the number of negative samples can vary with the actual difficulty level of the local region. Thus, the negative sample sampling quantity determination mechanism driven by regional difficulty information allows high-confidence background regions, target edge regions, strongly textured background regions, and local regions with a high risk of false detection to receive more negative sample supervision, reducing the burden on the limited negative sample budget for simple background regions.

[0101] In some embodiments, the number of negative samples for each local region can be determined based on at least one of region difficulty information, feature map scale parameters, and training phase parameters. This allows the negative sample budget to be adaptively allocated according to changes in region difficulty, scale differences, and training phase. Feature map scale parameters may include feature map stride, feature map height, feature map width, feature map resolution, or the number of candidate prediction locations at that feature map scale. Since different feature map scales are typically responsible for detecting targets of different sizes, their background distribution and false detection risk may also differ. Different upper limits, lower limits, or sampling quantity adjustment coefficients can be set for different feature map scales to maintain a balance in multi-scale negative sample supervision.

[0102] In this invention application, during the early stages of training, the prediction results of the object detection model are still unstable. A smaller and smoother number of negative samples can be set to ensure a more balanced sampling opportunity for each local region, avoiding excessive amplification of the impact of a few high-confidence noise backgrounds on the training process. In the later stages of training, the object detection model has already acquired a certain discriminative ability, allowing for the allocation of more negative samples to local regions with higher difficulty levels. This enables the model to further learn high-confidence false detection backgrounds, object edge backgrounds, and easily confused category backgrounds.

[0103] In some embodiments, a minimum and a maximum number of negative samples can be set. The minimum number of samples can be used to ensure that local regions have basic negative sample supervision opportunities, while the maximum number of samples can be used to limit a single local region from consuming too much negative sample budget. For example, for local regions with low difficulty, the number of negative samples can be set to zero or a small value; for local regions with high difficulty, the number of samples can be appropriately increased within the maximum number of samples limit. Thus, while maintaining spatial coverage, more of the negative sample budget can be allocated to high-value regions.

[0104] In some embodiments, the number of negative samples is no greater than the number of candidate negative samples within the corresponding local region. That is, when the number of negative samples calculated based on at least one of the region difficulty information, feature map scale parameters, and training phase parameters is greater than the actual number of candidate negative samples existing within the local region, the number of negative samples can be limited to the number of candidate negative samples within that local region; when there are no candidate negative samples within the local region, the number of negative samples corresponding to that local region can be zero. This ensures that the sampling process is executable and avoids situations where no candidate objects are available for sampling or where invalid locations are repeatedly selected.

[0105] In some embodiments, after determining the number of negative samples corresponding to each local region, candidate negative samples can be sampled within each local region according to the negative sample sampling probability and the number of negative samples. Specifically, for any local region, the negative sample sampling probability corresponding to each candidate negative sample within that local region can be used as the sampling basis, and the number of negative samples corresponding to that local region can be used as the sampling count to select a corresponding number of candidate negative samples from that local region. The negative samples obtained from sampling each local region can be merged to form a set of region-sampled negative samples.

[0106] In some embodiments, candidate negative samples within each local region can be sampled without replacement according to the negative sample sampling probability. Sampling without replacement can be understood at least as follows: once a candidate negative sample is selected within the same local region, it will not participate in subsequent sampling within that local region. By sampling without replacement, the same candidate negative sample can be avoided from being selected repeatedly in a single sampling process, resulting in a set of sampled negative samples containing more diverse background candidate locations, thus improving the training value and sample diversity of the sampled negative samples.

[0107] In some embodiments, when the number of negative samples corresponding to a certain local region is one, a candidate negative sample for a region can be selected according to the negative sample sampling probability of each candidate negative sample in that local region. When the number of negative samples corresponding to a certain local region is greater than one, multiple candidate negative samples for regions can be selected consecutively without replacement, and the selected candidate negative samples for regions can be removed from the candidate pool after each selection. In this way, the tendency for negative samples with high sampling probabilities to be selected is maintained, while avoiding the same candidate negative sample from repeatedly entering the set of negative samples for region sampling.

[0108] In some embodiments, if the number of candidate negative samples within a certain local region is less than or equal to the number of negative samples sampled for that local region, then all candidate negative samples within that local region can be determined as the region-sampled negative samples. If the number of candidate negative samples within a certain local region is greater than the number of negative samples sampled for that local region, then sampling without replacement can be performed based on the negative sample sampling probability to select candidate negative samples with greater training value. Thus, both cases of insufficient and sufficient candidate negative samples can be considered.

[0109] Optionally, quality filtering is performed on the negative samples in the region sampling negative sample set to remove sampled negative samples that meet the low-value sample condition. This includes: calculating the quality score corresponding to each negative sample in the region sampling negative sample set, wherein the quality score is determined based on the target confidence score and / or negative sample loss; identifying negative samples with quality scores lower than a preset quality threshold as low-value samples and removing the low-value samples from the region sampling negative sample set; wherein the preset quality threshold is a fixed threshold, or is adjusted according to at least one of the following: training stage, feature map scale, local region difficulty, and historical false detection statistics.

[0110] In some embodiments, a quality score can be calculated for each negative sample in the region-sampled negative sample set. The quality score can characterize the effectiveness or difficulty of the sampled negative samples in training the target detection model. A higher quality score indicates that the corresponding negative sample is more likely to be misclassified as a target by the model, or that it contributes more to the target confidence loss; a lower quality score indicates that the corresponding negative sample is more likely to be identified as background by the model, or that it provides a weaker training gradient. Therefore, the quality score can be used to filter the sampled negative samples in the region-sampled negative sample set.

[0111] In some embodiments, the quality score can be determined based on the target confidence score. The target confidence score can be obtained by activating the target confidence prediction value corresponding to the negative sample in the region sampling. For a negative sample with the true label as background, if its target confidence score is high, it indicates that the model still tends to believe that the candidate location contains the target, and the negative sample has a high risk of false detection and training value; if its target confidence score is low, it indicates that the model can easily identify the candidate location as background, and the negative sample is more likely to be a simple negative sample. Therefore, the target confidence score can be used as the quality score, or as an important component of the quality score.

[0112] In other embodiments, the quality score can also be determined based on the negative sample loss. The negative sample loss can be the target confidence loss calculated by setting the target confidence value of the negative samples sampled from the region as the background label. For example, when the target confidence loss uses binary cross-entropy loss, for negative samples whose true label is the background, the higher the target confidence prediction value output by the model, the larger the corresponding negative sample loss usually is, indicating that the negative sample is more difficult for the model to correctly identify as background. Therefore, the negative sample loss can be used as a quality score, making it easier to retain difficult negative samples with larger losses.

[0113] Therefore, the quality score can be determined based on at least one of the target confidence score and the negative sample loss. For example, the target confidence score can be directly used as the quality score; the negative sample loss can be directly used as the quality score; or the target confidence score and the negative sample loss can be weighted and fused to obtain the quality score. In these ways, the training value of sampled negative samples can be measured from two perspectives: prediction confidence and loss contribution, making the quality filtering results more consistent with the learning needs of difficult negative samples in the training process of object detection models.

[0114] In some embodiments, after obtaining the quality score corresponding to each negative sample in the region-sampled negative sample set, the quality score can be compared with a preset quality threshold. When the quality score of a sampled negative sample is lower than the preset quality threshold, the sampled negative sample can be identified as a low-value sample and removed from the region-sampled negative sample set. When the quality score of a sampled negative sample is greater than or equal to the preset quality threshold, the sampled negative sample can be considered to have certain training value and can be retained in the region-sampled negative sample set after quality filtering.

[0115] In some embodiments, low-value samples can be understood as simple background negative samples that contribute little to the training of the object detection model. For example, when the target confidence score of a sampled negative sample is very low, it indicates that the model has been able to stably identify the candidate location as background, and the negative sample provides little training information for further participation in the target confidence loss calculation; when the negative sample loss of a sampled negative sample is very low, it indicates that the difference between the prediction result corresponding to the negative sample and the background label is small, and its contribution to the loss is limited. For the aforementioned low-value samples, they can be removed from the region-sampled negative sample set to avoid the limited negative sample budget being consumed by simple background samples.

[0116] In some embodiments, the preset quality threshold can be a fixed threshold. For example, when the quality score uses a target confidence score, the preset quality threshold can be set to 0.01, 0.03, 0.05, 0.10, or other empirical thresholds; when the quality score uses negative sample loss, the preset quality threshold can be set to a loss threshold that matches the numerical range of the loss function. Fixed thresholds are simple to implement, facilitate consistent filtering standards across different training batches, and are suitable for scenarios where the training data distribution is relatively stable or the model training phase does not change significantly.

[0117] In some embodiments, the preset quality threshold can also be adaptively adjusted according to the training stage. In the early stages of training, when the prediction results of the object detection model are still unstable, a lower preset quality threshold can be set to retain more negative samples from different regions, thus avoiding premature removal of samples that may still have training value. In the later stages of training, when the object detection model has acquired a certain background discrimination capability, the preset quality threshold can be appropriately increased to more strictly filter low-value simple background samples. This avoids simple background samples occupying the limited negative sample budget, allowing the training process to focus more on high-confidence backgrounds and difficult negative samples.

[0118] In other embodiments, the preset quality threshold can also be adjusted according to the feature map scale. Different feature map scales typically correspond to detection tasks for targets of different sizes, and their number of candidate negative samples, target confidence distribution, and false detection risk may differ. For example, for high-resolution feature map scales, since there are many candidate prediction locations and dense background regions, a relatively high quality threshold can be used to reduce a large number of simple background negative samples from entering subsequent loss calculations; for low-resolution feature map scales, since there are fewer candidate prediction locations, a relatively low quality threshold can be used to avoid over-filtering leading to insufficient supervision of negative samples at this scale.

[0119] In some other embodiments, the preset quality threshold can be adjusted based on the difficulty information of the local region. For local regions with higher difficulty, a relatively low or moderate quality threshold can be set to retain more difficult background samples that may have training value; for local regions with lower difficulty, a relatively high quality threshold can be set to more strictly filter simple background samples. This allows the quality filtering strategy to be matched with the actual false detection risk of the local region.

[0120] In some embodiments, the preset quality threshold can also be adaptively adjusted based on historical false detection statistics. Historical false detection statistics may include the distribution of false detection locations, false detection categories, false detection feature map scales, false detection region types, or the frequency of false detections for different background categories obtained during training or validation. If a certain feature map scale, a certain local region type, or a certain category of related background has a high number of false detections in historical statistics, the quality threshold for the corresponding region or scale can be adjusted to retain more potentially valuable negative samples in that region. This allows the negative sample quality filtering mechanism to establish a feedback relationship with the model's actual false detection performance.

[0121] In some embodiments, after removing low-value samples from the region-sampled negative sample set, the number of removed negative samples can be recorded. This number can be used to determine the number of negative samples to be backfilled in the subsequent hard negative sample backfilling step. Thus, the quality filtering step not only removes low-value samples but also provides a quantitative basis for selecting high-hardness negative samples from unsampled candidate negative samples for backfilling, thereby maintaining a relatively stable number of negative samples ultimately used in the target confidence loss calculation.

[0122] In some embodiments, quality filtering can be performed independently within each training image, or separately within each batch according to the image dimension or feature map scale dimension. For example, a quality score can be calculated and threshold filtering performed by sampling a set of negative samples for each region of the training image; alternatively, quality thresholds can be set separately for different feature map scales and filtering can be performed accordingly. In this way, the quality filtering process can be adapted to the differences in negative sample distribution across different images, scales, and local regions.

[0123] Therefore, by performing quality filtering on the negative sample set of the region sampling, we can further remove simple background negative samples with low training value while maintaining the regional balanced sampling framework. This reduces the impact of invalid or inefficient gradients on the target detection model training process, provides a clear amount of compensation for subsequent difficult negative sample backfilling, and allows the removed low-value negative samples to be replaced by high-value difficult negative samples from the unsampled candidate negative samples, thereby improving the overall quality of the negative samples that finally participate in the target confidence loss calculation.

[0124] Optionally, from the unsampled candidate negative samples, difficult negative samples corresponding to the number of removed negative samples are selected for backfilling according to their difficulty scores, including: The negative sample set of the region sampling is excluded from the candidate negative sample set to obtain the backfill candidate negative sample set; the difficulty score corresponding to each backfill candidate negative sample in the backfill candidate negative sample set is calculated, and the difficulty score is determined according to at least one of the target confidence loss, target confidence prediction value, target confidence score, class confusion score, and class entropy; backfill negative samples corresponding to the number of removed negative samples are selected from the backfill candidate negative sample set in descending order of the difficulty score; when the number of candidates in the backfill candidate negative sample set is less than the number of removed negative samples, all candidate negative samples in the backfill candidate negative sample set are determined as the backfill negative samples; and / or, at least one of the following is set for the backfill negative samples: backfill ratio upper limit, feature map scale backfill upper limit, and local region backfill upper limit, to limit the concentration of the backfill negative samples in the local region or at the feature map scale.

[0125] And / or, obtaining the candidate negative sample set includes: The prediction results of the object detection model at at least one feature map scale are obtained. The prediction results include the target confidence prediction value, class prediction value, and bounding box prediction value corresponding to the candidate prediction position. Based on the real annotation information of the training image and the positive sample allocation algorithm, a set of positive samples is determined from the candidate prediction positions, and a foreground mask is generated based on the set of positive samples. According to the foreground mask, positive sample positions are excluded from the candidate prediction positions, thereby obtaining a set of candidate negative samples.

[0126] In this invention application, after quality filtering of the sampled negative sample set and counting the number of removed negative samples, difficult negative samples can be further selected from the unsampled candidate negative samples for backfilling. Since the quality filtering step removes sampled negative samples with low target confidence, small negative sample loss, or low training value, without compensation, the number of negative samples ultimately participating in the target confidence loss calculation may decrease. Therefore, based on the number of removed negative samples, a corresponding number of high-difficulty negative samples can be selected from the unsampled candidate negative samples for backfilling to ensure that the final negative sample set has both regional balance and high training value.

[0127] In some embodiments, the region-sampled negative sample set can be excluded from the candidate negative sample set to obtain a backfill candidate negative sample set. The candidate negative sample set can be understood at least as the set of background candidate positions obtained after excluding positive sample positions from the candidate prediction positions of the object detection model. The region-sampled negative sample set can be understood at least as the set of negative samples already selected through local region sampling. By excluding the region-sampled negative sample set from the candidate negative sample set, the same candidate negative sample can be avoided from being both a sampled negative sample and a backfill negative sample, thus ensuring that the backfilling process can supplement new difficult background samples from the candidate negative samples that have not yet been selected by region sampling.

[0128] In some embodiments, the backfill candidate negative sample set may include candidate negative samples that did not enter the regional sampling negative sample set. That is, for a candidate prediction location, if it belongs to the candidate negative sample set and was not selected in the aforementioned regional sampling process, it can be used as a backfill candidate negative sample. In this way, the source of supplementary difficult negative samples can be expanded, so that high-difficulty negative samples that were not selected due to regional budget constraints in the aforementioned regional sampling process still have the opportunity to enter the final negative sample set.

[0129] In some embodiments, a difficulty score can be calculated for each backfill candidate negative sample in the backfill candidate negative sample set. The difficulty score can be used to characterize the risk that the backfill candidate negative sample will be misclassified as a target or a certain target category by the target detection model. The higher the difficulty score, the more difficult it is for the corresponding backfill candidate negative sample to be correctly identified as background by the model, and the higher its training value; the lower the difficulty score, the easier it is for the corresponding backfill candidate negative sample to be identified as background by the model, and the lower its priority as a backfill negative sample.

[0130] In some embodiments, the difficulty score can be determined using the target confidence loss. The target confidence loss can be the loss value calculated when the target confidence value for backfilling candidate negative samples is set to the background label. For example, when the target confidence loss uses binary cross-entropy loss, for candidate negative samples with the true label as the background, the higher the predicted target confidence value output by the model, the larger the target confidence loss corresponding to that candidate negative sample is usually, indicating that the model is more likely to misclassify it as the target. Therefore, the target confidence loss can be used as the difficulty score to prioritize backfilling high-loss background negative samples.

[0131] In other embodiments, the difficulty score can be determined by the target confidence prediction value or the target confidence score. The target confidence prediction value can be the target confidence logical value output by the object detection model, and the target confidence score can be the value obtained after activating the target confidence prediction value. For backfilling candidate negative samples, the higher the target confidence prediction value or the target confidence score, the more likely the background candidate location is to be predicted by the model as containing the target. Therefore, the difficulty score can be determined based on the target confidence prediction value or the target confidence score, making it easier for high target confidence background samples to be selected as backfilling negative samples.

[0132] In some embodiments, the difficulty score can also be determined by a category confusion score. The category confusion score can be determined based on the category prediction value corresponding to the backfilled candidate negative sample, historical false positive statistics, a category confusion matrix, or prior knowledge of the business category. In applications such as industrial defect detection, some background textures, although not belonging to the real target, may be close to the defect category or other target categories in the category branch, easily leading to category false positives during the inference stage. By incorporating the category confusion score into the difficulty score, the probability of backfilling background candidate positions that are more similar to the target in terms of category can be increased.

[0133] In other embodiments, the difficulty score can also be determined using category entropy. Category entropy can be used to characterize the uncertainty of the model's classification of backfill candidate negative samples. When the category entropy is high, it indicates that the predicted distribution of the backfill candidate negative sample on the category branch is relatively scattered, and the model is uncertain about its category assignment. When category entropy is used in combination with the target confidence score, target confidence loss, or category confusion score, it can more comprehensively characterize the false detection risk of background candidate positions. Therefore, the backfilling priority of uncertain background samples or category boundary background samples can be increased based on category entropy.

[0134] Clearly, the difficulty score can be determined individually by one of the following: target confidence loss, target confidence predicted value, target confidence score, class confusion score, and class entropy; or it can be determined by a combination of these indicators. For example, the target confidence loss can be used as the primary difficulty indicator, with the class confusion score or class entropy used as an auxiliary modulation term; alternatively, the target confidence score, class confusion score, and class entropy can be weighted and fused to obtain a comprehensive difficulty score. In this way, the training value of backfilling candidate negative samples can be evaluated from two perspectives: the risk of misjudgment regarding the existence of the target and the risk of misjudgment regarding the class.

[0135] In some embodiments, after obtaining the difficulty score corresponding to each backfill candidate negative sample, the backfill candidate negative sample set can be sorted in descending order of difficulty score, and backfill negative samples corresponding to the number of removed negative samples can be selected from the sorting results. The number of removed negative samples can be the number of low-value negative samples removed from the negative sample set sampled from the region during the quality filtering step. By making the number of backfill negative samples correspond to the number of removed negative samples, the same or similar number of difficult negative samples can be added after removing low-value negative samples, so that the number of negative samples participating in the final target confidence loss calculation remains relatively stable.

[0136] In some embodiments, when the number of removed negative samples is zero, hard negative sample backfilling may not be performed, or the backfilled negative sample set may be set to an empty set. When the number of removed negative samples is greater than zero, several backfilled candidate negative samples with the highest difficulty scores can be selected from the backfilled candidate negative sample set based on this number. Thus, the hard negative sample backfilling process can only play a compensatory role when the quality filtering reduces the number of negative samples, avoiding the introduction of additional excessive negative samples that would change the scale of the training loss.

[0137] In some embodiments, when the number of candidates in the backfill candidate negative sample set is less than the number of removed negative samples, all candidate negative samples in the backfill candidate negative sample set can be determined as backfill negative samples. That is, when the number of candidate negative samples available for backfilling is insufficient to compensate for the number of removed negative samples, all available candidate negative samples can be used for compensation as much as possible, rather than forcibly constructing non-existent backfill samples. This ensures that the backfilling process conforms to the actual number of candidate negative samples and avoids invalid or duplicate positions from participating in the backfilling process.

[0138] Therefore, by using the hard negative sample backfilling method, after removing low-value sampled negative samples through quality filtering, high-difficulty background samples from unsampled candidate negative samples can be added in a timely manner, thereby improving the training value of the final negative sample set. By using the correspondence between backfilling quantities and the upper limit constraint of backfilling, the stability of the final negative sample quantity, spatial distribution, and multi-scale distribution can be maintained, thus forming an effective balance between regional balanced sampling and hard negative sample mining, and improving the target detection model's ability to suppress false detections of high-confidence backgrounds, easily confused category backgrounds, and target edge backgrounds.

[0139] In some embodiments, at least one of the following can be set for backfilling negative samples: backfilling ratio upper limit, feature map scale backfilling upper limit, and local region backfilling upper limit: Specifically: The upper limit on the backfill ratio can be used to limit the proportion of backfilled negative samples in the final negative sample set. For example, the number of backfilled negative samples can be limited to a certain proportion of the number of negative samples retained after quality filtering, or a certain proportion of the final negative sample set. By setting an upper limit on the backfill ratio, excessive backfilled negative samples can be avoided, which weakens the spatial coverage effect formed by the aforementioned balanced regional sampling. The upper limit on feature map scale backfill can be used to limit the number of backfilled negative samples from the same feature map scale. Since background candidate locations with high difficulty scores may be concentrated at a certain feature map scale, if backfilling is performed entirely according to the global difficulty ranking, it may lead to an excessive concentration of backfilled negative samples at a single scale, thus affecting the multi-scale negative sample supervision balance. By setting an upper limit on feature map scale backfill, a relatively reasonable distribution of backfilled negative samples can be maintained between different feature map scales. The upper limit on local region backfill can be used to limit the number of backfilled negative samples from the same local region. Since some complex texture regions, reflective regions, target edge regions, or regions near annotation boundaries may contain multiple high-difficulty candidate negative samples, without restrictions, the backfilled negative samples may be concentrated in a few local regions. By setting an upper limit for backfilling in local areas, it is possible to supplement high-difficulty negative samples while avoiding disrupting the spatial balance brought about by regional balanced sampling.

[0140] Clearly, the upper limit for backfill ratio, the upper limit for feature map scale backfill, and the upper limit for local region backfill can be used individually or in combination. For example, candidate negative samples can be first sorted according to their difficulty scores, and then candidate negative samples that satisfy the upper limits for backfill ratio, feature map scale backfill, and local region backfill can be selected sequentially. When a candidate negative sample causes the corresponding upper limit to be exceeded, that candidate negative sample can be skipped, and subsequent candidate negative samples in the sorting results can be examined. Thus, scale balance and spatial balance can be considered while prioritizing difficulty.

[0141] In some embodiments, obtaining a candidate negative sample set may include: obtaining the prediction results of an object detection model at at least one feature map scale, where the prediction results may come from one or more detection heads of the object detection model or from prediction outputs on multi-scale feature maps; determining a positive sample set from candidate prediction locations based on ground truth annotation information of training images and a positive sample allocation algorithm; and generating a foreground mask based on the positive sample set. Ground truth annotation information may include ground truth bounding boxes, ground truth class labels, and optional ignored region annotations. The positive sample allocation algorithm may determine which candidate prediction locations are used to perform the ground truth object prediction task based on the matching relationship between candidate prediction locations and ground truth objects. The foreground mask may be used to mark whether a candidate prediction location belongs to a positive sample location, wherein locations marked as foreground by the foreground mask can be used as locations in the positive sample set.

[0142] In some embodiments, positive sample locations can be excluded from candidate predicted locations based on a foreground mask, thereby obtaining a candidate negative sample set. Locations in the candidate predicted locations that are not marked as positive samples by the foreground mask can serve as the basic source of candidate negative samples. Furthermore, invalid filling locations, locations corresponding to ignored regions, or other locations unsuitable for participating in negative sample supervision can also be excluded, resulting in a valid candidate negative sample set. This candidate negative sample set can serve as the basic sample set for subsequent local region segmentation, region sampling, quality filtering, and hard negative sample backfilling.

[0143] Secondly, such as Figure 2 As shown, a negative sample sampling system, employing or using any one of the negative sample sampling methods described in the first aspect above, is applied to target detection model training, comprising: Module 100 is configured to acquire a set of candidate negative samples. The region segmentation module 200 is configured to divide the candidate prediction positions at each feature map scale into multiple local regions, and determine the regional candidate negative samples belonging to the candidate negative sample set within each local region. The sampling parameter determination module 300 is configured to construct the sampling probability of negative samples in each local region based on the target confidence prediction value corresponding to the candidate negative samples in the region, and determine the sampling quantity of negative samples corresponding to each local region. The region sampling module 400 is configured to sample candidate negative samples in each local region according to the negative sample sampling probability and the negative sample sampling quantity, thereby obtaining a set of region-sampled negative samples. The quality filtering module 500 is configured to perform quality filtering on negative samples in the negative sample set of the region, remove sampled negative samples that meet the low-value sample condition, and count the number of removed negative samples. The difficult backfilling module 600 is configured to select difficult negative samples corresponding to the number of removed negative samples from the unsampled candidate negative samples and backfill them according to the difficulty score, so as to obtain a backfilled negative sample set. The loss construction module 700 is configured to construct a target confidence loss for training the target detection model based on the positive sample set, the region sampling negative sample set retained after quality filtering, and the backfill negative sample set.

[0144] In some embodiments, the configurations of the acquisition module 100, the region division module 200, the sampling parameter determination module 300, the region sampling module 400, the quality filtering module 500, the difficult backfill module 600, and the loss construction module 700 can be found in the same or related technical content in the first aspect, and will not be repeated here. Accordingly, the relevant descriptions or contents in the first aspect regarding the acquisition of candidate negative sample sets, local region division, determination of regional candidate negative samples, construction of negative sample sampling probability, determination of negative sample sampling quantity, sampling without replacement within the region, quality filtering of low-value negative samples, difficult negative sample backfill, and construction of target confidence loss can all be applied to the negative sample sampling system of this aspect.

[0145] In some embodiments, the acquisition module 100, region segmentation module 200, sampling parameter determination module 300, region sampling module 400, quality filtering module 500, hard backfilling module 600, and loss construction module 700 can be independent software functional modules, or they can be integrated into the same object detection training program, loss function program, training script, deep learning framework component, or model training plugin. These modules can interact with each other through function calls, tensor transfer, state caching, training context objects, shared memory, message queues, or other software communication methods.

[0146] In the negative sample sampling system of this invention, the acquisition module 100 can acquire a set of candidate negative samples for sampling; the region division module 200 can divide the multi-scale candidate prediction location into multiple local regions; the sampling parameter determination module 300 can determine the negative sample sampling probability and the number of negative samples in the local region; the region sampling module 400 can obtain the set of negative samples sampled in the region; the quality filtering module 500 can remove low-value sampled negative samples; the hard backfilling module 600 can fill in high-hardness negative samples from the unsampled candidate negative samples; and the loss construction module 700 can construct the target confidence loss based on positive samples, retained negative samples, and backfilled negative samples. Through the cooperation between the above modules, similarly, negative sample sampling combining region balancing, hardness awareness, quality filtering, and hard backfilling can be achieved during the target detection model training process, thereby improving the efficiency of negative sample utilization and the ability of the target detection model to suppress false detections in complex backgrounds.

[0147] Optionally, such as Figure 3As shown, the acquisition module 100 includes: a prediction result acquisition unit 101 configured to acquire the prediction results of the target detection model at at least one feature map scale, the prediction results including the target confidence prediction value, category prediction value, and bounding box prediction value corresponding to the candidate prediction position; a positive sample determination unit 102 configured to determine a set of positive samples from the candidate prediction positions based on the real annotation information of the training image and the positive sample allocation algorithm, and generate a foreground mask based on the set of positive samples; and a candidate negative sample determination unit 103 configured to exclude positive sample positions from the candidate prediction positions according to the foreground mask to obtain a set of candidate negative samples.

[0148] In some embodiments, the configuration of each unit within the acquisition module 100 can refer to the same or related technical content in the first aspect, and will not be repeated here. That is to say, the relevant records or contents in the first aspect regarding the target detection model prediction results, candidate prediction positions, positive sample sets, positive sample allocation algorithms, foreground masks, positive sample position exclusion, invalid position exclusion, and candidate negative sample set acquisition can all be applied to the acquisition module 100 and the units constituting it in this aspect.

[0149] In this invention application, through the cooperation of the prediction result acquisition unit, the positive sample determination unit, and the candidate negative sample determination unit, the candidate prediction results of the target detection model on the multi-scale feature map can be obtained first, and the set of positive samples used to undertake the real target prediction task can be determined using real annotation information. Then, positive sample positions can be excluded from the candidate prediction positions based on the foreground mask to obtain the set of candidate negative samples used for subsequent region division, region sampling, quality filtering, and difficult backfilling. This allows the negative sample sampling system to independently process the candidate negative samples after the positive samples are determined without changing the positive sample allocation logic, thereby improving the compatibility of the negative sample sampling system with the existing target detection training framework.

[0150] Thirdly, this application provides a negative sample sampling device, including a memory and a processor connected in communication, wherein the memory is used to store a computer program, and the processor is used to read the computer program and execute the negative sample sampling method described in any of the first aspects above.

[0151] Those skilled in the art will understand that the negative sample sampling apparatus includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the negative sample sampling method described in any of the first aspects above.

[0152] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, perform any of the negative sample sampling methods described in the first aspect above.

[0153] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.

[0154] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0155] Wherein, if the modules / units integrated in the terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0156] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without creative effort.

[0157] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A negative sample sampling method applied to target detection model training, characterized in that, include: Obtain the candidate negative sample set; The candidate prediction locations at each feature map scale are divided into multiple local regions, and regional candidate negative samples belonging to the candidate negative sample set are determined within each local region. Based on the target confidence prediction value corresponding to the candidate negative samples in the region, the negative sample sampling probability in each local region is constructed, and the negative sample sampling quantity corresponding to each local region is determined; according to the negative sample sampling probability and the negative sample sampling quantity, the candidate negative samples in the region are sampled in each local region to obtain the region sampled negative sample set; The negative samples in the negative sample set of the region are quality filtered to remove the sampled negative samples that meet the low value sample condition, and the number of removed negative samples is counted; from the unsampled candidate negative samples, difficult negative samples corresponding to the number of removed negative samples are selected according to the difficulty score and backfilled to obtain the backfilled negative sample set. Based on the positive sample set, the negative sample set of the region sampled after quality filtering, and the backfilled negative sample set, a target confidence loss is constructed for training the target detection model.

2. The negative sample sampling method according to claim 1, characterized in that, The candidate prediction positions at each feature map scale are divided into multiple local regions, including: obtaining the feature map height, feature map width, and region size corresponding to each feature map scale; dividing the candidate prediction positions at the corresponding feature map scale into multiple local regions according to the region size; when the feature map height or feature map width cannot be divided by the region size, the candidate prediction positions at the feature map scale are padded, and the padded positions are marked as invalid positions; when determining the regional candidate negative samples within each local region, positive sample positions and invalid positions are excluded; the region size is a preset fixed size or the region size should be determined according to at least one of the step size of the corresponding feature map scale, feature map size, input image size, target scale statistics, and training phase parameters.

3. The negative sample sampling method according to claim 1, characterized in that, Based on the target confidence prediction value corresponding to the candidate negative samples in the region, the negative sample sampling probability in each local region is constructed, including: The target confidence prediction value corresponding to the candidate negative sample in the region is activated to obtain the target confidence score; based on the target confidence score and the difficulty adjustment parameter, the basic sampling weight of the candidate negative sample in the region is determined; the basic sampling weight of each candidate negative sample in the same local region is processed to obtain the negative sample sampling probability in the corresponding local region; wherein, the difficulty adjustment parameter changes dynamically with the training progress to make the negative sample sampling probability distribution in the early stage of training smoother and to make the negative sample sampling probability in the later stage of training more biased towards difficult negative samples with high target confidence.

4. The negative sample sampling method according to claim 3, characterized in that, The basic sampling weights of the candidate negative samples in each of the aforementioned local regions are processed to obtain the negative sample sampling probability in the corresponding local region, including: Based on the predicted class value corresponding to the candidate negative sample in the region, a class-aware weight is determined, which is determined according to at least one of the maximum class probability, class entropy, class confusion score, and class prior weight; based on the neighborhood relationship between the candidate negative sample and the positive sample, a positive sample neighborhood modulation weight is determined; based on the feature map scale corresponding to the candidate negative sample in the region, a scale modulation weight is determined. Based on the basic sampling weight, and at least one of the category-aware weight, the positive sample neighborhood modulation weight, and the scale modulation weight, the comprehensive sampling weight of the candidate negative sample in the region is obtained, and the sampling probability of the negative sample is obtained based on the comprehensive sampling weight.

5. The negative sample sampling method according to claim 1, characterized in that, Determining the number of negative samples for each of the aforementioned local regions includes: The system collects at least one of the following regional difficulty information within each local region: the number of regional candidate negative samples, the number of high-confidence candidate negative samples, the mean target confidence score, and the maximum target confidence score. Based on at least one of the regional difficulty information, feature map scale parameters, and training phase parameters, the system determines the number of negative samples to be sampled for each local region. The number of negative samples sampled is not greater than the number of regional candidate negative samples within the corresponding local region. And / or, sampling candidate negative samples in each of the local regions according to the negative sample sampling probability and the negative sample sampling quantity includes: sampling candidate negative samples in each of the local regions without replacement according to the negative sample sampling probability.

6. The negative sample sampling method according to claim 1, characterized in that, Quality filtering is performed on negative samples in the negative sample set of the region, removing sampled negative samples that meet the low-value sample criteria, including: Calculate the quality score corresponding to each negative sample in the negative sample set of the region sampling, the quality score being determined based on the target confidence score and / or negative sample loss; identify negative samples with quality scores lower than a preset quality threshold as low-value samples and remove the low-value samples from the negative sample set of the region sampling; wherein, the preset quality threshold is a fixed threshold, or is adjusted based on at least one of the following: training stage, feature map scale, local region difficulty, and historical false detection statistics.

7. The negative sample sampling method according to claim 1, characterized in that, From the candidate negative samples that have never been sampled, difficult negative samples corresponding to the number of removed negative samples are selected for backfilling according to their difficulty scores, including: The negative sample set of the region is excluded from the candidate negative sample set to obtain the backfill candidate negative sample set; the difficulty score corresponding to each backfill candidate negative sample in the backfill candidate negative sample set is calculated, and the difficulty score is determined according to at least one of the target confidence loss, target confidence prediction value, target confidence score, class confusion score, and class entropy; backfill negative samples corresponding to the number of removed negative samples are selected from the backfill candidate negative sample set in descending order of the difficulty score; when the number of candidates in the backfill candidate negative sample set is less than the number of removed negative samples, all candidate negative samples in the backfill candidate negative sample set are determined as the backfill negative samples; And / or, obtaining the candidate negative sample set includes: The prediction results of the object detection model at at least one feature map scale are obtained. The prediction results include the target confidence prediction value, class prediction value, and bounding box prediction value corresponding to the candidate prediction position. Based on the real annotation information of the training image and the positive sample allocation algorithm, a set of positive samples is determined from the candidate prediction positions, and a foreground mask is generated based on the set of positive samples. According to the foreground mask, positive sample positions are excluded from the candidate prediction positions, thereby obtaining a set of candidate negative samples.

8. The negative sample sampling method according to claim 1, characterized in that, At least one of the following is set for the backfilled negative samples: backfill ratio upper limit, feature map scale backfill upper limit, and local area backfill upper limit, in order to limit the concentration of the backfilled negative samples in the local area or at the feature map scale.

9. A negative sample sampling system, applied to target detection model training, characterized in that, include: The module retrieves a set of candidate negative samples. The region segmentation module is configured to divide the candidate prediction location at each feature map scale into multiple local regions, and determine the regional candidate negative samples belonging to the candidate negative sample set within each local region. The sampling parameter determination module is configured to construct the sampling probability of negative samples in each local region based on the target confidence prediction value corresponding to the candidate negative samples in the region, and determine the sampling quantity of negative samples corresponding to each local region. The region sampling module is configured to sample candidate negative samples in each local region according to the negative sample sampling probability and the negative sample sampling quantity, thereby obtaining a set of region-sampled negative samples; The quality filtering module is configured to perform quality filtering on negative samples in the negative sample set of the region, remove sampled negative samples that meet the low-value sample condition, and count the number of removed negative samples. The difficult backfilling module is configured to select difficult negative samples from the unsampled candidate negative samples according to the difficulty score, which corresponds to the number of removed negative samples, and backfill them to obtain a backfilled negative sample set. The loss construction module is configured to construct a target confidence loss for training the target detection model based on the positive sample set, the region sampling negative sample set retained after quality filtering, and the backfilled negative sample set.

10. A negative sample sampling device, characterized in that, Including processor and memory, The memory stores a computer program, and when the processor executes the computer program, it performs the negative sample sampling method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the negative sample sampling method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Positive and negative sample candidate box-based dynamic selection ship detection method and system

    CN116758429A

  • Landslide disaster negative sample optimization method based on improved frequency ratio

    CN120123777A