Target detection online learning dynamic sample selection method and system, computer equipment and storage medium
By employing an online learning method for object detection that combines multi-dimensional feature evaluation and dynamic weight adjustment, the limitations of sample selection strategies and inappropriate resource utilization are addressed, thereby improving the model's adaptability and detection accuracy.
Patent Information
- Application Number
- CN202511539105.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing object detection models suffer from problems in online learning, such as a single dimension for sample value evaluation, fixed feature weights, improper resource utilization, and insufficient annotation quality. These issues lead to low model update efficiency and insufficient adaptability.
Monte Carlo Dropout sampling is used to obtain sample uncertainty, construct multi-dimensional feature vectors, dynamically adjust weights through attention mechanism, and combine with computing power constraints to perform efficient sample selection. An incremental learning strategy is used to update the model.
It achieves more comprehensive sample value assessment, dynamic adaptive weight adjustment, and efficient resource utilization, thereby improving the model's robustness and generalization ability in complex environments.
Smart Images

Figure CN120997492A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and artificial intelligence, and specifically relates to an online learning dynamic sample selection method, system, computer equipment, and storage medium for object detection. Background Technology
[0002] Object detection is one of the core tasks in the field of computer vision, and it is widely used in scenarios such as autonomous driving, intelligent monitoring, and industrial quality inspection. As the dynamics and complexity of application scenarios increase, object detection models are required to have online learning capabilities. That is, they should be able to continuously optimize their performance using newly collected data without interrupting service, in order to adapt to dynamic factors such as changes in environment, lighting, and the addition of new object categories.
[0003] Traditional online learning methods often rely on a single, fixed evaluation metric when selecting samples for model updates, such as the model's classification confidence score for the samples alone. Samples with low confidence scores are considered to contain information unknown to the model and are therefore prioritized. However, this strategy has significant drawbacks: 1. Limited Dimensions in Sample Value Assessment: Only considering classification confidence scores while ignoring other crucial information. For example, a sample may have high classification confidence but poor bounding box regression accuracy; such samples are still valuable for model optimization. Furthermore, the rarity of the scenario represented by the sample and the quality of the sample annotation should also be important dimensions for evaluating its value.
[0004] 2. Fixed Feature Weights: In some methods that integrate multiple features for sample evaluation, the weights of each feature are usually pre-set or empirically fixed. However, the focus of performance improvement may differ at different stages of the model's lifecycle or when facing different application scenarios. For example, in the early stages of a model, it may be more necessary to improve the overall recognition rate; while in the mature stage of a model, the focus may be more on solving detection problems in specific challenging scenarios. Fixed feature weights cannot be adaptively adjusted, limiting the flexibility and effectiveness of sample selection strategies.
[0005] 3. Ignoring computational resource constraints: Online learning is typically conducted on resource-constrained edge devices or servers requiring high timeliness. Most existing methods do not fully consider the actual constraints of computational resources, which may lead to update timeouts due to selecting too many samples, or insufficient model performance improvement due to selecting too few samples, failing to achieve the optimal balance between resource utilization and model benefits.
[0006] 4. Insufficient consideration of annotation quality: Online learning samples may come from automatically labeled (weakly labeled) or manually labeled, resulting in inconsistent quality. Low-quality or erroneous annotations can seriously mislead model updates and reduce model reliability. Current technologies generally lack quantitative evaluation mechanisms for sample annotation quality.
[0007] Therefore, how to design an online learning sample strategy that can comprehensively and dynamically evaluate sample value and make efficient selections under resource constraints is a technical problem that urgently needs to be solved in the field of object detection. Summary of the Invention
[0008] The purpose of this invention is to provide a method, system, computer device, and storage medium for online learning dynamic sample selection for target detection. The aim is to achieve more efficient, robust, and adaptive online updates of the target detection model by comprehensively evaluating sample value, dynamically adjusting evaluation weights, and being aware of computing power limitations.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: An online learning dynamic sample selection method for object detection includes the following steps: S1: The sample to be screened is input into the target detection model, and the classification uncertainty (U) of the sample to be screened is obtained through Monte Carlo Dropout sampling. c ) and positioning uncertainty (U l ); S2: Construct a multi-dimensional feature vector for the samples to be screened. The multi-dimensional feature vector includes at least the classification uncertainty (U). c ), positioning uncertainty (U l Knowledge gap matching degree (Im), annotation confidence (Sa), and scenario scarcity (Sc); S3: Construct a dynamic weight learning network based on the attention mechanism, and calculate the weight of each feature in the multi-dimensional feature vector by combining the performance index of the validation set after the online update of the object detection model; S4: Based on the feature weights calculated in step S3, perform multi-dimensional value scoring on the samples to be screened, and select the Top-K samples with the highest value scores to form the optimal sample subset, where K is the number of samples determined based on the online learning computing power limit. S5: Input the optimal sample subset into the target detection model to complete the online update.
[0010] Furthermore, in step S1: Classification uncertainty (U c The entropy is calculated using the category probability distribution entropy, and the formula is as follows: , where p k Let be the probability of the k-th category, and K be the total number of categories in the object detection model.
[0011] Positioning uncertainty (U l The detection box coordinates are calculated using the mean of the sampling variance of the detection box coordinates. These coordinates include the center coordinates (x, y) and the width and height (w, h). The calculation formula is as follows: , These represent the sampling variances for the corresponding coordinates.
[0012] Furthermore, in step S2: The knowledge gap matching degree (Im) is calculated using mutual information theory, with the formula Im = (X;M), where X is the feature of the sample to be screened, specifically the high-dimensional global feature vector output by the backbone network of the object detection model (ResNet-50 or YOLOv8-backbone); M is the feature distribution of the historical training samples of the object detection model, fitted by a Gaussian mixture model (GMM), and the number of components of the GMM is determined by the Bayesian information criterion (BIC).
[0013] The label confidence level (Sa) is determined as follows: if the sample to be screened is manually labeled, Sa = 1; if it is weakly labeled, Sa = 1 - label error rate. The label error rate is calculated through the historical label validation set. Specifically, N weakly labeled samples are selected as the validation set (N≥100), and the number of incorrectly labeled samples n is determined by manual review. The label error rate = n / N.
[0014] Scene scarcity (Sc) is calculated by dividing the historical proportion of the scenes to which the samples to be screened belong by the inverse, using the following formula: , This represents the number of samples for this scene in the historical training set of the object detection model.
[0015] Furthermore, the structure of the dynamic weight learning network in step S3 includes a feature mapping layer, an attention calculation layer, and a weight output layer, as detailed below: 1. Feature Mapping Layer: This layer maps the eigenvalues (Ui, U ...) of a multi-dimensional feature vector. c U l After being standardized to the [0, 1] interval, the features are mapped to a 128-dimensional feature vector through a fully connected layer, and the activation function is ReLU. 2. Attention Calculation Layer: Inputs performance metrics of the target detection model's validation set (including mAP improvement rate ΔmAP, loss reduction rate ΔLoss, and generalization error change ΔE). gen The attention score is obtained by performing a dot product operation with the 128-dimensional feature vector output by the feature mapping layer. , where F perf For the performance metric feature vector (dimension 128), F feat Output of the feature mapping layer; 3. Weighted Output Layer: This layer calculates the attention score A. f Normalization is performed to obtain the weights of each feature. , satisfying ∑w=1.
[0016] Furthermore, the training method for the dynamic weight learning network is as follows: The optimization objective is to maximize the validation set mAP of the object detection model after the next round of updates. The Adam optimizer is used for end-to-end training with a learning rate of 1e-4 and 30 iterations. The training samples in each round are paired data of "historical multi-dimensional feature vectors + corresponding model performance indicators", with a sample size of no less than 200 sets.
[0017] Furthermore, in step S4: The multi-dimensional value score is calculated as follows: for each sample i to be screened, a value score is calculated. ,in Let be the standardized feature values of sample i; The logic for determining the sample size K is as follows: T max M is the maximum allowed duration for a single online update. GPU T represents the GPU memory capacity. single M represents the training time for a single sample. single For single sample memory usage, This indicates rounding down to the nearest integer.
[0018] Furthermore, the online update strategy for the object detection model in step S5: An incremental learning strategy is adopted: only the fully connected layers and the detection head layer of the model are updated, the backbone network layers are frozen, the learning rate is set to 1 / 10 of the initial training learning rate, and the number of iterations is 5-10 rounds.
[0019] An online learning dynamic sample selection system for object detection includes the following modules: Uncertainty Sampling Module: Used to perform step S1 above and output the classification uncertainty (U) of the sample to be screened. c ) and positioning uncertainty (U l ); Multi-dimensional feature construction module: used to perform step S2 above and generate multi-dimensional feature vectors; Attention weight learning module: used to perform step S3 above, outputting the real-time weights of each feature based on the attention mechanism and model validation set performance metrics; Value scoring and sampling module: used to perform step S4 above, calculate multi-dimensional value scores for samples and select the optimal sample subset; Model update and feedback module: This module is used to execute step S5 above, realize online model updates, and feed back the updated validation set performance metrics to the attention weight learning module.
[0020] This application also proposes a computer device, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program stored in the memory to implement the method described in any of the above.
[0021] This application also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the steps of the method described in any one of the above-described embodiments.
[0022] This computer-readable storage medium, by executing the computer program instructions, stores the collected data on the user terminal itself, preventing data leakage. The collected data is added to the training process, and new data is integrated into a training batch and then added to the global model for training. This increases the amount of data and prevents overfitting, making the dynamically learned model more generalizable.
[0023] Beneficial effects Compared with the prior art, the present invention has the following beneficial effects: (1) Comprehensive sample value assessment: By integrating the features of five dimensions, namely classification uncertainty, positioning uncertainty, knowledge gap, annotation quality and scenario scarcity, this invention can more accurately and comprehensively assess the real value of samples for model updates than single index methods, avoiding the opportunity cost caused by one-sided selection.
[0024] (2) Dynamically adaptive weight strategy: A dynamic weight learning network based on attention mechanism is introduced, which enables the sample selection strategy to automatically adjust the weights of each feature according to the real-time performance feedback of the model (such as changes in indicators such as mAP and Loss). This allows the method to intelligently adapt to the needs of different scenarios and model lifecycles, and achieve "on-demand sampling".
[0025] (3) Resource-aware and efficient sampling: This invention uses computing power limitation as the key parameter to determine the number of samples, and dynamically calculates the optimal number of samples K through a formulaic approach. This ensures that model updates can make full use of available resources to achieve maximum performance improvement, while avoiding update failures or timeouts due to resource overload, thus achieving the best balance between model benefits and resource costs.
[0026] (4) Improved robustness and generalization ability: By introducing label confidence (Sa), the negative impact of low-quality or mislabeled samples on the model is effectively reduced. At the same time, by using the scene scarcity (Sc) feature, scene samples that the model has not seen or rarely see are selected first, which helps the model learn richer environmental features and significantly improves its generalization ability and robustness in complex and ever-changing real environments.
[0027] (5) Efficient incremental update strategy: Combining the incremental learning method of freezing the backbone network, the computational overhead and time cost in the online update process are greatly reduced, making high-frequency model iteration possible and enhancing the real-time adaptability of the model. Attached Figure Description
[0028] Figure 1 This is a flowchart of the method provided by the present invention; Figure 2 This is a system module diagram provided by the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to specific embodiments.
[0030] This embodiment takes target detection (such as pedestrian and vehicle recognition) in intelligent monitoring scenarios as an example to illustrate the specific implementation process of the present invention. The target detection model used is Faster R-CNN based on ResNet-50 backbone network. The hardware environment is a single NVIDIA Tesla V100 GPU (32GB of video memory), and the software environment is the PyTorch 1.13 framework.
[0031] This embodiment provides an online learning dynamic sample selection method for object detection. The method execution flow is as follows: Figure 1 The specific execution flow of this invention is as follows: 1. Data preparation and initial model training ① Dataset construction: Historical labeled data in intelligent monitoring scenarios were selected as the initial training set, containing 30,000 images, covering 10 scenarios such as "daytime streets", "nighttime parking lots", and "rainy day intersections". The labeled categories include pedestrians, cars, buses and trucks, and the labeled format is the COCO standard format (including detection box coordinates and category labels).
[0032] ② Initial Model Training: The Faster R-CNN model was trained using the dataset described above, with a ResNet-50 backbone network. The initial learning rate was set to 1e-3, and the SGD optimizer (momentum 0.9) was used for 120 iterations. After training, the model achieved a mean average precision (mAP) of 78.5% on the validation set (5000 images), serving as the benchmark model for online learning.
[0033] 2. Input of samples to be screened and calculation of uncertainty (corresponding to step S1) ① Source of samples to be screened: 20,000 unlabeled images generated in real time by the intelligent monitoring system, covering the newly added scenarios of "foggy highways" and "congested road sections", as candidate samples for online learning.
[0034] ② Monte Carlo Dropout Sampling: Enable Dropout (dropout rate = 0.5) in the backbone network and detection head layer of the object detection model, perform 10 independent forward propagations for each sample to be screened, and obtain 10 sets of class probability distributions and detection box coordinates.
[0035] ③ Classification uncertainty (U c Calculate: For the prediction of the "pedestrian" category in a certain sample, the probability distribution obtained from 10 samplings is [p1, p2, ..., p...]. 10 ], where the class probability of the k-th sample is p k =[0.82, 0.05, 0.10, 0.03] (corresponding to pedestrians, cars, buses, and trucks), then the classification uncertainty of this sample is:
[0036] ④ Positioning uncertainty (U) l Calculate: For 10 samples of the "pedestrian" detection box, the variance of the center coordinate x. =0.002, the variance of y =0.003, variance of width w =0.001, variance of height h =0.002, then the positioning uncertainty is:
[0037] 3. Construction of multi-dimensional feature vectors (corresponding to step S2) ① Knowledge gap matching degree (Im): Extracting high-dimensional features X from the samples to be screened: Outputting a 512-dimensional global feature vector through the avgpool layer of ResNet-50.
[0038] Fitting the historical feature distribution M: A Gaussian mixture model (GMM) is trained using the features of 30,000 samples from the initial training set, and the number of GMM components is determined to be 8 using the Bayesian information criterion (BIC).
[0039] Computing mutual information The joint probability distribution of X and M was estimated using the Monte Carlo method, and the Im value of a certain "foggy highway" sample was 0.32 (the lower the value, the greater the difference between the sample and the historical distribution, and the greater the knowledge gap).
[0040] ② Label the confidence level (Sa): Of the samples to be screened, 10,000 were manually labeled (Sa=1), and 10,000 were weakly labeled (automatically labeled using a traditional object detection model).
[0041] Weak labeling error rate calculation: 100 weakly labeled samples are randomly selected as the validation set. After manual review, 8 samples are found to be incorrectly labeled (e.g., a truck is incorrectly labeled as a bus). Therefore, the labeling error rate = 8 / 100 = 0.08. Hence, Sa = 1 - 0.08 = 0.92 for the weakly labeled samples.
[0042] ③Scene scarcity (Sc): The number of samples for the "foggy highway" scene in the historical training set is 0, while the number of samples for the "daytime street" scene is 5000.
[0043] The Sc=1 / 0=1 for the "Foggy Highway" sample (the default value for scarce scenarios is 1), and the Sc=1 / 5000=0.0002 for the "Daytime Street" sample.
[0044] 4. Dynamic weight learning network training and weight calculation (corresponding to step S3) ①Network structure and training data: The dynamic weight learning network takes a 5-dimensional feature vector (U) as input. c U l (Im, Sa, Sc), the output is a 5-dimensional weight vector.
[0045] The training data consists of paired data of "historical features + performance metrics": 1000 sets of sample features from the past 5 model updates are collected, along with the corresponding updated mAP improvement rate ∆mAP, loss reduction rate ∆Loss, and generalization error change ΔE. gen .
[0046] Training configuration: With the goal of "maximizing mAP in the next round", the Adam optimizer (learning rate 1e-4) is used, with 30 rounds of iteration and a batch size of 32 per round.
[0047] Example of weight calculation: Feature mapping layer: Maps the standardized features of a sample [0.68, 0.002, 0.32, 0.92, 1] to a 128-dimensional vector F through a fully connected layer. feat .
[0048] Attention computation layer: Input the current model's validation set performance metrics ∆mAP = 0.05, ∆Loss = -0.12, ΔE gen = -0.03, mapped to a 128-dimensional vector F feat Calculate the attention score by calculating the dot product.
[0049] Weighted output layer: After normalization, the weights are obtained as w = [0.25, 0.15, 0.20, 0.10, 0.30] (∑w = 1).
[0050] 5. Sample selection and model update (corresponding to steps S4 and S5) ① Multi-dimensional value rating (S): For a weakly labeled sample of "foggy highway", its standardized feature is U c = 0.68, U l = 0.002, Im = 0.32, Sa = 0.92, Sc = 1, then the value score is: Si = 0.25×0.68 + 0.15×0.002 + 0.20×0.32 + 0.10×0.92 + 0.30×1=0.17 + 0.0003 + 0.064 + 0.092 + 0.3 = 0.6263 ② Determine the sample size K: Given T max = 60 minutes (3600 seconds), M GPU = 32GB, T single = 0.8 seconds / sample, M single =1024MB (1GB) / sample, then:
[0051] Since there are only 20,000 samples to be screened, the actual number of samples screened is Top-20,000.
[0052] ③ Online model updates: An incremental learning strategy was adopted: the ResNet-50 backbone network was frozen, and only the fully connected layers and the detector head were updated. The learning rate was set to 1 / 10 (1e-4) of the initial learning rate, and the process was repeated for 8 rounds.
[0053] Performance improvements after the update: validation set mAP increased to 82.3% (an improvement of 3.8%), and detection accuracy in the "foggy highway" scene increased from 65% to 88%.
[0054] 6. System Module Interaction Flow Reference Figure 2 This invention can be implemented as an online learning dynamic sample selection system for target detection, comprising: Uncertainty sampling module: Receives samples to be screened and outputs U through Monte Carlo Dropout. c and U l The data is then transmitted to the multi-dimensional feature construction module.
[0055] Multi-dimensional feature construction module: Combines historical data to calculate Im, Sa, and Sc, generates a complete feature vector, and sends it to the attention weight learning module.
[0056] Attention weight learning module: Calculates feature weights based on the model's current performance metrics (such as ΔmAP) and feeds them back to the value scoring and sampling module.
[0057] Value scoring and sampling module: Calculates the value score of the sample, selects the Top-K samples, and submits them to the model update and feedback module.
[0058] Model update and feedback module: After completing the incremental update of the model, the new performance metrics are fed back to the attention weight learning module to achieve dynamic iteration.
[0059] This application also proposes a computer device, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program stored in the memory to implement the above-described target detection online learning dynamic sample selection method.
[0060] This application also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the steps of the above-described target detection online learning dynamic sample selection method.
[0061] This computer-readable storage medium, by executing the computer program instructions, stores the collected data on the user terminal itself, preventing data leakage. The collected data is added to the training process, and new data is integrated into a training batch and then added to the global model for training. This increases the amount of data and prevents overfitting, making the dynamically learned model more generalizable.
[0062] This embodiment demonstrates that the present invention, through multi-dimensional feature evaluation and dynamic weight adjustment, can select the most valuable samples under computational constraints, significantly improving the adaptability and detection accuracy of the target detection model in dynamic scenes.
[0063] In summary, this invention provides a systematic, intelligent, and resource-aware online learning sample selection scheme for target detection, effectively solving many pain points in the prior art and providing strong technical support for building a high-performance target detection system that can continuously adapt to dynamic environments.
[0064] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for online learning dynamic sample selection for object detection, characterized in that, Includes the following steps: S1: The sample to be screened is input into the target detection model, and the classification uncertainty (U) of the sample to be screened is obtained through Monte Carlo Dropout sampling. c ) and positioning uncertainty (U l ); The classification uncertainty (U) c The entropy is calculated using the category probability distribution entropy, and the formula is as follows: , where p k Let K be the probability of the k-th class, and K be the total number of classes in the object detection model. The positioning uncertainty (U) l The detection box coordinates are calculated using the mean of the sampling variance of the detection box coordinates. These coordinates include the center coordinates (x, y) and the width and height (w, h). The calculation formula is as follows: , These represent the sampling variances of the corresponding coordinates; S2: Construct a multi-dimensional feature vector for the samples to be screened. The multi-dimensional feature vector includes at least the classification uncertainty (U). c ), positioning uncertainty (U l Knowledge gap matching degree (Im), annotation confidence (Sa), and scenario scarcity (Sc); The knowledge gap matching degree (Im) is calculated using information theory mutual information, with the formula Im = (X;M), where X is the feature of the sample to be screened, specifically the high-dimensional global feature vector output by the backbone network of the object detection model (ResNet-50 or YOLOv8-backbone); M is the feature distribution of the historical training samples of the object detection model, fitted by a Gaussian mixture model (GMM), and the number of components of the GMM is determined by the Bayesian information criterion (BIC). The label confidence level (Sa) is determined as follows: if the sample to be screened is manually labeled, Sa = 1; if it is weakly labeled, Sa = 1 - label error rate. The label error rate is calculated through the historical label validation set. Specifically, N weakly labeled samples are selected as the validation set (N ≥ 100), and the number of incorrectly labeled samples n is determined by manual review. The label error is n / N. The scenario scarcity (Sc) is calculated by the reciprocal of the historical proportion of the scenarios to which the samples to be screened belong, using the following formula: , This represents the number of samples for this scene in the historical training set of the object detection model; S3: Construct a dynamic weight learning network based on the attention mechanism, and calculate the weight of each feature in the multi-dimensional feature vector by combining the performance index of the validation set after the online update of the object detection model; S4: Based on the feature weights calculated in step S3, perform multi-dimensional value scoring on the samples to be screened, and select the Top-K samples with the highest value scores to form the optimal sample subset, where K is the number of samples determined based on the online learning computing power limit. S5: Input the optimal sample subset into the target detection model to complete the online update.
2. The method according to claim 1, characterized in that, The structure of the dynamic weight learning network described in step S3 includes a feature mapping layer, an attention calculation layer, and a weight output layer. The execution flow is as follows: Feature mapping layer: This layer maps the eigenvalues (Ui, U ...) of a multi-dimensional feature vector. c U l After being standardized to the [0, 1] interval, the features (Im, Sa, Sc) are mapped to a 128-dimensional feature vector through a fully connected layer, and the activation function is ReLU; Attention computation layer: Inputs the validation set performance metrics of the object detection model (including mAP improvement rate ΔmAP, loss reduction rate ΔLoss, and generalization error change ΔE). gen The attention score is obtained by performing a dot product operation with the 128-dimensional feature vector output by the feature mapping layer. , where F perf For the performance metric feature vector (dimension 128), F feat Output of the feature mapping layer; Weighted output layer: This layer calculates the attention score A. f Normalization is performed to obtain the weights of each feature. , satisfying ∑w=1.
3. The method according to claim 2, characterized in that, The training method of the dynamic weight learning network is as follows: with the optimization objective of maximizing the validation set mAP after the next round of object detection model update, the Adam optimizer is used for end-to-end training, the learning rate is set to 1e-4, and the number of iterations is 30 rounds; the training samples in each round are paired data of "historical multi-dimensional feature vectors + corresponding model performance indicators", and the sample size is no less than 200 sets.
4. The method according to claim 1, characterized in that, The multi-dimensional value score calculation method in step S4 is as follows: For each sample i to be screened, calculate the value score. ; in Let be the standardized feature value of sample i.
5. The method according to claim 1, characterized in that, The logic for determining the sample quantity K in step S4 is as follows: ; T max M is the maximum allowed duration for a single online update. GPU T represents the GPU memory capacity. single M represents the training time for a single sample. single For single sample memory usage, This indicates rounding down to the nearest integer.
6. The method according to claim 1, characterized in that, The online update of the target detection model in step S5 adopts an incremental learning strategy: only the fully connected layers and the detection head layer of the model are updated, the backbone network layers are frozen, the learning rate is set to 1 / 10 of the initial training learning rate, and the number of iterations is 5-10 rounds.
7. A target detection online learning dynamic sample selection system, characterized in that, include: Uncertainty sampling module: used to execute step S1 in claim 1 and output the classification uncertainty (U) of the sample to be screened. c ) and positioning uncertainty (U l ); Multi-dimensional feature construction module: used to execute step S2 in claim 1 to generate a multi-dimensional feature vector; Attention weight learning module: used to execute step S3 in claim 1, and output the real-time weights of each feature based on the attention mechanism and the performance index of the model validation set; Value scoring and sampling module: used to perform step S4 in claim 1, calculate multi-dimensional value scores of samples and select the optimal sample subset; Model update and feedback module: used to execute step S5 in claim 1, realize online update of the model, and feed back the updated validation set performance index to the attention weight learning module.
8. A computer device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the steps of the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Visual target detection method and device based on vector quantization and uncertainty perception
CN120279314A
Lightweight multi-target instance segmentation method and system for inspection robot
CN120298692A
Container small target semi-supervised identification method and system
CN120495733A
Cited By
Deep learning core training sample selection and weight calibration method and system for image classification
CN121505311A
A deep learning core training sample selection and weight calibration method and system for image classification
CN121505311B