Visual target detection method and device based on vector quantization and uncertainty perception

Through vector quantization and uncertainty perception visual object detection methods, the problem that traditional methods cannot adapt to new target categories in the medical environment is solved, efficient detection of unknown classes and continuous accuracy maintenance of known classes is achieved, and open-world object detection is suitable for medical robots.

CN120279314APending Publication Date: 2025-07-08TIANJIN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510349535.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional object detection methods cannot adapt to the ever-changing new target categories in medical environments, and there are problems of known category offsets, pseudo-label misjudgment, and catastrophic forgetting in incremental learning, resulting in a degradation in model detection performance in complex medical environments.

Method used

The label allocation strategy based on vector quantization is adopted to map continuous high-dimensional features to discrete low-dimensional semantic space, decouple the known general prospect mode, and generate more accurate unknown class pseudo-labels by quantizing the positioning uncertainty and classification uncertainty of the model output, thereby improving detection performance.

Benefits of technology

While maintaining the detection accuracy of known classes, it significantly improves the performance of unknown classes, supports the continuous evolution and lifelong learning of medical robots, and does not require global retraining, and adapts to knowledge updates in complex medical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279314A_ABST
    Figure CN120279314A_ABST
Patent Text Reader

Abstract

The invention discloses a visual target detection method and device based on vector quantization uncertainty perception, and the method comprises the steps: collecting image data in an open scene as original data, marking the collected original data according to the category, and taking the marked original data as an initial task data set; training a target detection model on the initial task data set, wherein the target detection model comprises a target detection module, a vector quantization module and an uncertainty perception label distribution module; a new category of interest is screened from the explored unknown objects, and image data is collected and marked to serve as a new task data set; meanwhile, selecting a part of samples from the old task data set as a playback sample set; and finely adjusting the target detection model on the new task training set and the old task playback set to realize continuous expansion and evolution of visual target detection. The device comprises a processor and a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of open-world perception in computer vision, and particularly relates to a visual object detection method and device based on vector quantization and uncertainty perception, which are particularly applicable to object recognition in a medical robot vision system, can handle the discovery and dynamic evolution of new categories such as medical devices, and achieve continuously expanding visual object detection. Background Art

[0002] The object detection task realizes the perception of objects in an image through the joint modeling of bounding box localization and class classification, and is widely applied in fields such as medical robots and autonomous driving. In the field of medical robots, object detection not only concerns improving the diagnostic accuracy, but also plays a crucial role in aspects such as surgical navigation. However, the dynamic openness of the medical environment makes traditional object detection methods face many challenges. There are a wide variety of medical objects and they are constantly changing, and the continuous update of medical devices leads to the inability of traditional methods to meet new requirements. To solve this problem, an open-world object detection task applicable to the medical environment has been proposed. Different from traditional closed-world detection methods, open-world object detection aims to construct a detection framework that can dynamically evolve and can identify and adapt to new object categories in a changing environment. Current research attempts to construct an open-scene model through a generalization discrimination index based on the features of known categories, combined with a pseudo-label generation strategy. However, the performance improvement of these methods is still limited and cannot meet the requirements of complex environments in practical applications.

[0003] Traditional object detection methods usually assume that object categories are known during the training phase and can be learned from a limited dataset. However, in a medical environment, unknown objects constantly emerge, making traditional methods face difficulties. First, the generalization discrimination metrics optimized based on known categories often have distribution biases, causing the decision boundary of the model to favor known categories and unable to effectively identify new medical objects. Second, the unreliability of the pseudo-label generation mechanism can lead to background misjudgment and misclassification of known categories. In a medical environment, pseudo-labels usually cannot provide sufficient accuracy and are prone to mislabeling the background or irrelevant regions as known categories, affecting the overall accuracy of the model. The conflict between new and old knowledge in the incremental learning process is also a major challenge. As new knowledge is continuously introduced, the model may encounter "catastrophic forgetting", that is, losing the ability to recognize old categories when learning new categories, which is particularly serious in the object detection of medical robots and may lead to a decline in the recognition ability of historical diseases or surgical instruments. In addition, the conflict between new and old medical knowledge will also cause cross-contamination of features, further affecting the performance of the model. Only after these problems are effectively solved can the medical robot object detection system have the ability to adapt to complex medical environments. Therefore, constructing a medical robot object detection system that can dynamically learn, adapt to new objects, and maintain high accuracy and high robustness has become an important direction for current technological development. Summary of the Invention

[0004] The present invention provides a visual object detection method and device based on vector quantization and uncertainty perception. Through semantic feature space modeling based on vector quantization, the present invention maps continuous high-dimensional features to a discrete low-dimensional semantic space, decouples known general foreground patterns, thereby improving the recall performance for medical devices; at the same time, designs an uncertainty-aware label assignment strategy, generates more accurate pseudo-labels for unknown classes by quantifying the localization uncertainty and classification uncertainty of the model output, and improves the detection performance for medical devices, as described in detail below:

[0005] In a first aspect, a visual object detection method based on vector quantization and uncertainty perception, the method includes:

[0006] Collect image data in an open scenario as raw data, and label the collected raw data according to categories as an initial task dataset;

[0007] Train an object detection model on the initial task dataset, the object detection model includes: an object detection module, a vector quantization module, and an uncertainty-aware label assignment module;

[0008] Screen interesting new categories from the discovered unknown objects, collect image data and label it as a new task dataset; at the same time, select a part of the samples from the old task dataset as a replay sample set;

[0009] Fine-tune the object detection model on the new task training set and the old task replay set to achieve the continuous expansion and evolution of visual object detection.

[0010] Among them, the object detection module includes: a first feature extractor, a first encoder, a first decoder, and a task head. Based on the object detection module, known-class foregrounds are generated to complete the positioning and classification tasks of known-class objects.

[0011] Among them, the vector quantization module includes: a second encoder, a discrete codebook, a second decoder, a teacher model, and a foreground classification head structure. Based on the vector quantization module, the feature reconstruction task and the foreground-background binary classification task are optimized.

[0012] Among them, the uncertainty-aware label assignment module is composed of the original prediction output of the model, the variance of the positioning reference point offset, the entropy of the classification feature distribution, and the prediction output of the classification head of the vector quantization module.

[0013] Among them, the foreground classification head is:

[0014] Extract the ROI features on F according to the candidate region (x, y, w, h) vq Extract ROI features And generate the binary classification score s = Sigmoid(W f cls f roi ) through linear projection, and the decoder Reconstruct the teacher features.

[0015] Among them, the uncertainty-aware label assignment module calibrates the confidence of unknown classes using classification uncertainty and positioning uncertainty.

[0016] The classification uncertainty w obj First, calculate the binary classification entropy of each known-class prediction of the model, and then calculate the mean of the binary classification entropies of all known classes;

[0017] The positioning uncertainty w loc Then calculate the variance of the multi-layer prediction output of the first decoder of the model;

[0018] The classification uncertainty is used to calibrate the classification confidence of unknown classes, and the positioning uncertainty is used to calibrate the positioning confidence of unknown classes to obtain the calibrated confidence of unknown classes

[0019] Select the top-k background predictions with the highest confidence of unknown classes as the pseudo-labels of unknown classes to participate in the training.

[0020] Second aspect: A visual target detection device based on vector quantization uncertainty perception, the device comprising: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to cause the device to execute the method according to any one of the first aspect.

[0021] Third aspect: A computer-readable storage medium storing a computer program, the computer program comprising program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of the first aspect.

[0022] The beneficial effects of the technical solution provided by the present invention are as follows:

[0023] 1. The present invention proposes a visual target detection network based on vector quantization, optimizes the discrete codebook through the feature reconstruction task, maps the continuous depth features to the low-dimensional structured codebook space, overcomes the defect that the classification index in the traditional method is overly biased towards known classes, and while maintaining the detection accuracy of known classes, significantly improves the recall performance of unknown classes and the cross-scene generalization ability of the model.

[0024] 2. The present invention designs an uncertainty-aware label assignment strategy, jointly measures the localization uncertainty and classification uncertainty, dynamically corrects the confidence of unknown classes, generates more reliable pseudo-labels for unknown classes, alleviates the defect that the model ignores unknown class instances mixed in the background in the open scenario, and improves the detection performance; and feasibility experiments are verified on the open-world task and incremental learning task datasets composed of PASCAL VOC2007 and MSCOCO2017.

[0025] 3. The present invention supports the continuous evolution and lifelong learning of medical robots, can complete knowledge update without global retraining, is conducive to promoting the sustainable evolution of intelligent medical devices in real scenarios; and improves the detection performance of medical devices. Description of the Drawings

[0026] Figure 1 It is a flowchart of a visual target detection method based on vector quantization uncertainty perception proposed by the present invention.

[0027] Figure 2 It is a schematic diagram of a semantic feature space modeling module based on vector quantization proposed by the present invention.

[0028] Figure 3 It is a schematic diagram of an uncertainty-aware label assignment module proposed by the present invention. Detailed Embodiments

[0029] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below.

[0030] To solve the technical problems existing in the background art, an embodiment of the present invention proposes an object detection network based on vector quantization and uncertainty perception. The network includes: a vector quantization module and an uncertainty perception module. Among them, the vector quantization module is used to screen potential unknown-class objects, and the uncertainty perception module is used to calibrate model predictions and generate reliable unknown-class pseudo-labels. By combining the two, it effectively overcomes the problem of detecting unknown-class objects faced by traditional object detection networks in open scenarios, and provides an effective solution for sustainable model evolution in dynamic medical environments.

[0031] As Figure 1 shown, a visual object detection method based on vector quantization uncertainty perception consists of three components, and the functions of the content of each part are as follows:

[0032] I. Object Detection Module

[0033] In the embodiment of the present invention, Deformable DETR is used as the object detection module, which includes: a first feature extractor, a first encoder, a first decoder, and a task head. Among them, the feature extractor is ResNet50, both the first encoder and the first decoder are 6-layer Deformable Attention Transformers, and the task head includes a classification head and a localization head, both of which are linear layers.

[0034] The object detection module is responsible for generating known-class foregrounds and completing the localization and classification tasks of known-class objects.

[0035] II. Vector Quantization Module

[0036] In the embodiment of the present invention, a vector quantization module is constructed using a second encoder, a discrete codebook, a second decoder, a teacher model, and a foreground classification head structure. The second encoder includes: a cascaded second feature extractor and a feature mapping layer. The second feature extractor is shared with the baseline detector and is responsible for extracting high-dimensional depth features of the image; the feature mapping layer consists of a two-layer perceptron and is responsible for mapping high-dimensional features to a low-dimensional space. The discrete codebook is a set of prototype vectors updated by exponential moving average, and discretization is achieved by replacing low-dimensional continuous features through nearest neighbor queries. The second decoder consists of a single-layer Vision Transformer and is responsible for reconstructing the low-dimensional feature map to a high-dimensional space.

[0037] The vector quantization module is optimized through a feature reconstruction task and a foreground-background binary classification task. The feature reconstruction task calculates the error between the features reconstructed by the second decoder and the features of the teacher model; for the foreground-background binary classification task, first, ROI Align is used to extract regional features from the vector quantization feature map, and then the foreground confidence is obtained through a linear classification head, and the binary classification loss is calculated with the known-class annotation.

[0038] The vector quantization module discretizes high-dimensional continuous depth features into a series of fine-grained discrete semantic concepts, constructs a more general semantic feature space, alleviates the shift of the foreground classification head distribution towards known classes, and thus improves the unknown class detection performance.

[0039] III. Uncertainty-Aware Label Assignment Module

[0040] In the embodiments of the present invention, the localization uncertainty and classification uncertainty of the object detection model are quantified to obtain more reliable pseudo-labels for unknown classes. Among them, the localization uncertainty uses the variance of the multi-layer localization output, and the classification uncertainty uses the mean of the binary classification entropy of each known class. The two cooperate to calibrate the confidence of unknown classes and improve the reliability of pseudo-labels.

[0041] In a real medical environment, visual object detection needs to consider the perception of unknown classes in an open scene. This task can be formally described as:

[0042] Given an input image and its known class annotation information where is a predefined class set, is a set of target bounding boxes. The model needs to output a dual prediction including the known class detection result and the unknown class detection result .

[0043] Furthermore, the vector quantization module specifically includes:

[0044] Map the image I through the encoder to the continuous depth feature space, and the quantization layer realizes the discrete mapping between the feature and the codebook . The whole process can be expressed as

[0045] In the subsequent process, the foreground classification head extracts the ROI feature vq on F according to the candidate region (x, y, w, h) and generates a binary classification score s = Sigmoid(W f cls f roi ) through linear projection to participate in the subsequent discrimination task. The decoder reconstructs the teacher feature to make the low-dimensional discrete space have generalizable semantic relevance.

[0046] Furthermore, the uncertainty-aware module specifically includes:

[0047] To discover potential unknown-class objects from the background instances predicted by the model, it is first necessary to construct a discriminant criterion for unknown classes. In the embodiments of the present invention, the foreground confidence of the vector quantization module is used as the basic classification confidence s for unknown classes. obj , and the model localization prediction b i =(x i , y i , w i , h i ) is used to calculate the maximum intersection over union with all selective search boxes as the basic localization confidence for unknown classes Finally, by comprehensively considering the classification confidence and the localization confidence for unknown classes, the confidence s for unknown classes is obtained unk = s obj × s loc .

[0048] Subsequently, the classification uncertainty and the localization uncertainty are used to calibrate the confidence for unknown classes. For the classification uncertainty w obj , first calculate the binary cross-entropy of the predictions for each known class of the model, and then calculate the mean of the binary cross-entropies of all known classes. For the localization uncertainty w loc , calculate the variance of the multi-layer prediction outputs of the model decoder. Finally, the classification uncertainty is used to calibrate the classification confidence for unknown classes, and the localization uncertainty is used to calibrate the localization confidence for unknown classes, obtaining the calibrated confidence for unknown classes Select the top-k background predictions with the highest confidence for unknown classes as the pseudo-labels for unknown classes to participate in the training.

[0049] During the continuous evolution process, according to the newly discovered classes or new requirements in the actual deployment environment, collect labeled data, and retain a small number of old task samples for knowledge replay to achieve the continuous evolution of the object detection model.

[0050] Embodiment 1

[0051] The embodiments of the present invention provide a visual object detection method based on vector quantization and uncertainty perception, and the method includes the following steps:

[0052] 101: Construct a visual object detection framework responsible for the classification and localization tasks of known classes:

[0053] Use Deformable DETR as the baseline detector. During the training process, the training data only includes the labeled information of known classes, and the classification head of the detector is "known classes + 1", including all known classes and an unknown class.

[0054] 102: Construct a vector quantization module to model the discrete semantic feature space and the forward discriminant index:

[0055] The continuous depth features are mapped to a low - dimensional discrete semantic space through a vector quantization module. The vector quantization module realizes the reconstruction of the feature space by optimizing the discrete codebook, enabling the features of unknown classes to be adaptively embedded in adjacent semantic regions during the training process while retaining the discriminative ability between foreground and background.

[0056] 103: Construct a discriminant criterion for unknown class pseudo - labels to discover potential unknown class instances in the background during the training process:

[0057] To screen potential unknown class instances from the detector predictions, a basic unknown class confidence is constructed using the foreground confidence of the vector quantization module and the degree of position overlap between the model prediction and the auxiliary bounding box.

[0058] 104: Design an uncertainty - aware dynamic label assignment mechanism to obtain more reliable unknown class pseudo - labels;

[0059] Based on the unknown class confidence, the unknown class confidence is calibrated using the quantized classification uncertainty and localization uncertainty to obtain more accurate unknown class pseudo - labels.

[0060] 105: Implement progressive model optimization and deployment applications for open - world medical scenarios;

[0061] In the test environment, the unknown class instances detected by the model are stored, and then it is manually judged whether these instances belong to new classes, and the bounding box positions of the new class instances are corrected. When the number of discovered new classes accumulates to more than 20, the model is extended and trained on these new class data to gradually improve the model's detection ability for new classes. Finally, the detection framework can adapt to the continuous perception requirements of unknown targets in complex scenarios such as medical robots.

[0062] In summary, the embodiment of the present invention constructs a more generalizable foreground discriminant index through the vector quantization module to alleviate the deviation of foreground discrimination towards known classes; and calibrates the model prediction through the uncertainty - aware module to obtain a more reliable unknown class discriminant criterion. Relevant experiments verify the robustness and practicality of the proposed scheme in an open and dynamic environment.

[0063] Embodiment 2

[0064] The following further introduces the scheme in Embodiment 1 in combination with specific examples and calculation formulas, as detailed in the following description:

[0065] I. Data Preparation

[0066] The embodiment of the present invention is applied to the open - world object detection task. The embodiment of the present invention verifies the effectiveness of the proposed method on an open - world object detection dataset composed of PASCAL VOC2007 and MS COCO2017.

[0067] The PASCAL VOC2007 dataset is a classic benchmark dataset for image object detection and segmentation, consisting of 9,963 images covering 20 common object categories. The data mainly comes from publicly available images on the Internet, covering various daily scenarios such as people, animals, and vehicles. The resolution of each image is approximately 500×375 pixels and is stored in JPEG format. This dataset provides detailed bounding box annotations and segmentation mask annotations, with a total of 24,640 annotation instances, averaging about 2.5 instances per image. According to the standard division, the training and validation set uses 5,011 images, and the test set contains 4,952 images.

[0068] The MS COCO2017 dataset is a large-scale multi-task visual dataset containing more than 330,000 images covering 80 fine-grained object categories. The data sources are extensive, covering natural scenes, indoor and outdoor environments, and various daily life situations. The image resolutions are diverse, up to 640×480 pixels at most, and are all stored in JPEG format. This dataset provides rich annotation information, including instance-level bounding boxes, pixel-level segmentation masks, and keypoint annotations, with a total number of annotations exceeding 2,500,000, and an average of 7.7 annotation instances per image. The standard division uses 118,287 training sets, 5,000 validation sets, and 40,670 test sets. The COCO dataset has become the mainstream evaluation benchmark for current tasks such as object detection and instance segmentation due to its high annotation density and strong scene diversity.

[0069] On this basis, there are two common dataset divisions for open-world object detection tasks. One is to use the categories and training samples of PASCAL VOC2007 as the training set for the first task, and evenly divide the remaining 60 categories in MS COCO into the latter three tasks; the other is to directly divide the 80 categories in the MS COCO2017 dataset evenly into four tasks. Each of the two task divisions has its own challenges: the former has the problem of additional cross-domain generalization, while the latter needs to handle more potential unknown instances in the background.

[0070] II. Visual Object Detection Network Structure Based on Vector Quantization and Uncertainty Awareness

[0071] The visual object detection network structure based on vector quantization and uncertainty awareness in the embodiments of the present invention, as Figure 1 shown, the embodiments of the present invention include three components: a baseline detector, a vector quantization module, and an uncertainty-aware label assignment module.

[0072] In the embodiment of the present invention, the baseline detector uses Deformable DETR as the backbone network architecture, which consists of a feature extractor, an encoder, a decoder, and a task head. Among them, the feature extractor is ResNet50, both the encoder and the decoder are 6-layer Deformable Attention Transformers, and in the task head, both the classification head and the regression head are linear layers.

[0073] In the embodiment of the present invention, the vector quantization module consists of an encoder, a discrete codebook, a decoder, a teacher model, and a foreground classification head, as Figure 2 shown; among them, in addition to sharing the feature extractor with the baseline detector, the encoder is connected with a two-layer perceptron after the 3rd, 4th, and 5th layers of the feature extractor to map the dimension of the feature map to the dimension of the discrete codebook; the discrete codebook is a variable with a fixed length that can be updated and is used to replace the original continuous depth features according to the principle of the nearest distance; the decoder is a single-layer Vision Transformer; the teacher model is a Vision Transformer model pre-trained by DINO; the foreground classification head is a linear layer.

[0074] The uncertainty-aware label assignment module in the embodiment of the present invention is as Figure 3 shown, which uses the classification uncertainty predicted by the model to calibrate the classification confidence of unknown classes and uses the localization uncertainty to calibrate the localization confidence of unknown classes to obtain a more reliable discriminant criterion for pseudo-labels of unknown classes.

[0075] III. Evaluation Metrics and Protocols

[0076] To verify the performance of the object detector in an open dynamic scenario, from the perspective of the mastery of known knowledge and the exploration of potential unknown knowledge, the present invention uses the mean average precision of known classes (mAP@50) and the recall rate of unknown classes (U-Recall@50) to evaluate the method performance. Specifically:

[0077]

[0078] Among them, is the set of known classes, is the set of unknown classes, TP j is the number of true positives, FN j is the number of false negatives, and P k (r) represents the precision corresponding to class k at recall rate r.

[0079] IV. Usage Details of the Model

[0080] 1. Data augmentation: Due to limited computing resources, the embodiments of the present invention adopt the methods of random horizontal flipping, random scaling, and cropping to improve data diversity. Specifically, the image is horizontally flipped with a probability of 50%, and then one of the following data augmentation strategies is selected with a probability of 50%: Strategy 1, randomly scale to a certain target size; Strategy 2, randomly scale to a certain short-side size, randomly crop the image to a smaller size (between 384 and 600), and then randomly scale again.

[0081] 2. Model optimization: In the training process of the embodiments of the present invention, the batch size is 2, the AdamW optimization algorithm is used, the initial learning rate is set to 2×10 -4 , and the learning rate of the feature extractor is set to 2×10 -5 , the weight decay rate is 1×10 -4 . After 40 epochs, the learning rate decays, and then continue to train for 10 more epochs.

[0082] 3. Hyperparameter settings: The embodiments of the present invention use 100 object queries, the codebook size is set to 1024×32, and the teacher model dimension is 768, corresponding to the ViTBase model. The pseudo labels select the top-5 background instances with the highest confidence of the unknown class as the unknown class supervision signals to participate in the training.

[0083] 4. Loss settings: The loss in the embodiments of the present invention consists of four parts: the classification loss of object detection, using binary cross-entropy loss; the regression loss of object detection, using a combination of GIOU loss and L1 loss; the foreground-background binary classification loss, using binary cross-entropy loss; and the feature reconstruction loss, using a combination of cosine loss and L1.

[0084] The above several losses are weighted and summed in the above order at a ratio of 2:2:5:1:0.7:0.3 to obtain the final loss.

[0085] The embodiments of the present invention all have the following three key creative points:

[0086] I. A semantic feature space modeling method based on vector quantization is proposed

[0087] Technical effect: Through the vector quantization mapping of the discrete semantic space, this method reconstructs the continuous depth features into a more generalizable structured semantic grid, effectively decoupling the strong correlation between the foreground general pattern and the known classes. This design breaks through the classification index bias problem caused by traditional closed-set optimization, and still has a reasonable embedding ability for unknown class samples beyond the training distribution in an open scenario, laying a reliable feature discrimination space for label assignment.

[0088] II. An uncertainty-aware label assignment strategy is proposed

[0089] Technical effect: By quantifying the classification uncertainty and localization uncertainty, this strategy calibrates the basic unknown class discrimination index to obtain more accurate pseudo-labels for unknown classes. In the open dynamic visual object detection scenario, this strategy significantly improves the recall rate and localization accuracy of unknown class instances, and realizes the refined collaborative perception of known classes and unknown classes. Relevant experimental verifications have been carried out on the open-world object detection dataset and the incremental dataset.

[0090] In summary, the embodiment of the present invention proposes a vector quantization feature modeling and dynamic label assignment framework for open-world object detection: by modeling the discrete semantic space, it solves the problem of known class feature deviation; through the uncertainty-aware dynamic threshold mechanism, it alleviates the contradiction between background misjudgment and known class misclassification.

[0091] Embodiment 3

[0092] The method proposed in the embodiment of the present invention has been experimented on the two common open-world object detection task partitions and compared with the past open-world object detection methods.

[0093] For the task sets divided based on PASCAL VOC2007 and MS COCO2017, the experimental results are shown in Table 1: in the main experimental indicators reflecting the open scenario, the known class accuracy (CK mAP) and unknown class recall (U-Recall) of Task 1 are better than the past methods; at the same time, in the main experimental indicators reflecting the dynamic scenario, the new class accuracy (CK mAP) and old class accuracy (PK mAP) of Task 4 also have certain advantages. Therefore, it is proved that the method in the embodiment of the present invention can better utilize the existing data to model the open dynamic scenario, thus achieving better detection effects.

[0094] Table 1

[0095]

[0096]

[0097] For the task sets divided based on MS COCO2017, the experimental results are shown in Table 2: in all indicators, the method in the embodiment of the present invention is better than the past methods.

[0098] Table 2

[0099]

[0100] Embodiment 4

[0101] The embodiments of the present invention can be extended to incremental object detection. The method proposed in the embodiments of the present invention is compared with multiple incremental object detection methods on a common evaluation benchmark set for incremental object detection. The incremental object detection task uses the partition of PASCAL VOC 2007 as the evaluation benchmark set. The 20 categories of samples are divided into two tasks: old categories and new categories, and are divided into three settings: "10+10", "15+5", and "19+1" according to the ratio of new categories to old categories.

[0102] The experimental results under the "10+10" setting are shown in Table 3. In terms of the old category mAP index that measures the retention of old knowledge, the method proposed in the embodiments of the present invention is superior to past methods, and there is also a certain advantage in the final performance. Therefore, it is proved that the method in the embodiments of the present invention can maintain old knowledge while learning new knowledge, so as to realize the continuous expansion of the object detection system in dynamic scenarios.

[0103] Table 3

[0104]

[0105]

[0106] The experimental results under the "15+5" and "19+1" settings are shown in Table 4 and Table 5 respectively. These two settings with unbalanced new and old categories pose a greater test on the model's ability to protect old knowledge. The method in the embodiments of the present invention is superior to past methods in maintaining old category knowledge and final detection accuracy under the "15+5" setting, and also has a certain advantage under the "19+1" setting. Therefore, it is proved that the method in the embodiments of the present invention can better protect old knowledge, so as to achieve better detection performance.

[0107] Table 4

[0108] Method Year Old class mAP New class mAP Final mAP ILOD 2017 68.3 58.4 65.8 Faster ILOD 2020 71.6 56.9 67.9 ORE-EUBI 2021 71.8 58.7 68.5 OW-DETR 2022 72.2 59.8 69.4 Open World DETR 2022 74.7 56.9 70.2 PROB 2023 73.2 60.8 70.1 Ours - 75.4 59.3 71.4

[0109] Table 5

[0110]

[0111]

[0112] The experimental results of the embodiments of the present invention are shown in Table 6. This result shows the test performance of four variants of this method on the open-world object detection dataset. All variants are trained on the task 1 training set and tested on the test set. In different experiments, the training steps, other parameters, and evaluation metrics are the same. As shown in Table 6, this method achieves better results than other variants, and at the same time verifies that the vector quantization module and the uncertainty-aware label assignment module proposed in the embodiments of the present invention can effectively improve the performance of the object detector in open dynamic scenarios.

[0113] Table 6

[0114] Method U-Recall CK mAP Deformable DETR 5.7 59.2 + Vector quantization module 23.0 59.8 + Uncertainty-aware label assignment module 19.8 60.9 This method 22.1 61.3

[0115] Example 5

[0116] A visual target detection device based on vector quantization and uncertainty perception, the device includes: a processor and a memory, program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to make the device execute the following method steps in Example 1:

[0117] Collect image data in an open scene as raw data, and label the collected raw data according to categories as the initial task dataset;

[0118] Train a target detection model on the initial task dataset, and the target detection model includes: a target detection module, a vector quantization module, and an uncertainty perception label assignment module;

[0119] Screen out new categories of interest from the discovered unknown objects, collect image data and label it as the new task dataset; at the same time, select a part of the samples from the old task dataset as the replay sample set;

[0120] Fine-tune the target detection model on the new task training set and the old task replay set to achieve continuous expansion and evolution of visual target detection.

[0121] Among them, the target detection module includes: a first feature extractor, a first encoder, a first decoder and a task head, and generates known class foregrounds based on the target detection module to complete the positioning and classification tasks of known class objects.

[0122] Among them, the vector quantization module includes: a second encoder, a discrete codebook, a second decoder, a teacher model and a foreground classification head structure, and optimizes the feature reconstruction task and the foreground-background binary classification task based on the vector quantization module.

[0123] Among them, the uncertainty perception label assignment module consists of the original prediction output of the model, the offset variance of the positioning reference point, the entropy of the classification feature distribution, and the prediction output of the classification head of the vector quantization module.

[0124] Among them, the foreground classification head is:

[0125] Extract ROI features on F according to the candidate region (x, y, w, h) vq and generate a binary classification score s = Sigmoid(W through linear projection f cls ), and the decoder roi reconstructs the teacher features. ​

[0126] Among them, the uncertainty-aware label assignment module calibrates the confidence of unknown classes using classification uncertainty and localization uncertainty.

[0127] Classification uncertainty w obj First, calculate the binary classification entropy of the predictions of each known class of the model, and then calculate the average value of the binary classification entropies of all known classes.

[0128] Localization uncertainty w loc Then calculate the variance of the multi-layer prediction outputs of the first decoder of the model.

[0129] Classification uncertainty is used to calibrate the classification confidence of unknown classes, and localization uncertainty is used to calibrate the localization confidence of unknown classes, obtaining the calibrated confidence of unknown classes.

[0130] Select the top-k background predictions with the highest confidence of unknown classes as pseudo-labels of unknown classes to participate in training.

[0131] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be elaborated herein.

[0132] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as a computer, a single-chip microcomputer, a microcontroller, etc. Specifically, in implementation, the embodiments of the present invention do not limit the execution subject, and it is selected according to the needs in actual applications.

[0133] Data signals are transmitted between the memory and the processor through a bus, and the embodiments of the present invention will not be elaborated herein.

[0134] Based on the same inventive concept, the embodiments of the present invention also provide a computer-readable storage medium. The storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the method steps in the above embodiments.

[0135] The computer-readable storage medium includes, but is not limited to, flash memory, hard disk, solid-state drive, etc.

[0136] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the method description in the embodiments, and the embodiments of the present invention will not be elaborated herein.

[0137] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.

[0138] A computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions may be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that incorporates one or more available media. The available medium may be a magnetic medium or a semiconductor medium, etc.

[0139] In the embodiments of the present invention, unless otherwise specified for the models of each device, the models of other devices are not limited, and any device that can perform the above functions is acceptable.

[0140] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0141] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A visual object detection method based on vector quantization uncertainty perception, characterized in that The method includes: Collecting image data in an open scenario as raw data, and annotating the collected raw data according to categories as an initial task dataset; Training an object detection model on the initial task dataset, where the object detection model includes: an object detection module, a vector quantization module, and an uncertainty-aware label assignment module; Screening interesting new categories from the discovered unknown objects, collecting image data and annotating it as a new task dataset; at the same time, selecting a part of the samples from the old task dataset as a replay sample set; Fine-tuning the object detection model on the new task training set and the old task replay set to achieve continuous expansion and evolution of visual object detection.

2. The visual target detection method based on vector quantization uncertainty perception according to claim 1, characterized in that, The object detection module includes: a first feature extractor, a first encoder, a first decoder, and a task head, generating known-class foregrounds based on the object detection module to complete the localization and classification tasks of known-class objects.

3. A visual target detection method based on vector quantization uncertainty perception according to claim 1, characterized in that, The vector quantization module includes: a second encoder, a discrete codebook, a second decoder, a teacher model, and a foreground classification head structure, optimizing the feature reconstruction task and the foreground-background binary classification task based on the vector quantization module.

4. A visual target detection method based on vector quantization uncertainty perception according to claim 1, characterized in that The uncertainty-aware label assignment module is composed of the original prediction output of the model, the variance of the localization reference point offset, the entropy of the classification feature distribution, and the prediction output of the classification head of the vector quantization module.

5. A visual target detection method based on vector quantization uncertainty perception according to claim 3, characterized in that, The foreground classification head is: Extract the ROI feature on F according to the candidate region (x, y, w, h) vq Extract the ROI feature And generate the binary classification score s = Sigmoid(W through linear projection cls f roi ), and the decoder reconstructs the teacher feature.

6. A visual target detection method based on vector quantization uncertainty perception according to claim 4, characterized in that The uncertainty-aware label assignment module calibrates the confidence of unknown classes using classification uncertainty and localization uncertainty, Classification uncertainty w obj First, calculate the binary classification entropy predicted by the model for each known class, and then calculate the mean of the binary classification entropies of all known classes; Positioning uncertainty w loc Then calculate the variance of the multi-layer prediction output of the first decoder of the computing model; Classification uncertainty is used to calibrate the classification confidence of unknown classes, and localization uncertainty is used to calibrate the localization confidence of unknown classes, so as to obtain the calibrated confidence of unknown classes Selecting the top-k background predictions with the highest confidence of unknown classes as unknown-class pseudo-labels to participate in training.

7. A visual target detection device based on vector quantization uncertainty perception, characterized in that, The device includes: a processor and a memory, where program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, where the computer program includes program instructions, and when the program instructions are executed by the processor, the processor executes the method described in any one of claims 1-6.

Citation Information

Cited By

  • Target detection online learning dynamic sample selection method and system, computer equipment and storage medium

    CN120997492A

  • An object detection online learning dynamic sample selection method, system, computer device and storage medium

    CN120997492B