Unsupervised domain adaptive open world target detection method

Through the unsupervised domain adaptation method, preliminary pseudo-labels are generated and the images are segmented into source domain and target domain, and the threshold is dynamically adjusted to improve the generalization ability of the model, solving the problems of unknown object detection bias and knowledge forgetting in the prior art, and achieving more efficient training and detection performance.

CN120032111APending Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510132463.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing open-world object detection methods rely on the supervisory information of known objects when detecting unknown objects, resulting in the model having bias to detect unknown objects and may forget the knowledge learned before during incremental learning.

Method used

Using an unsupervised domain adaptation method, the images are divided into source domain and target domain by generating preliminary pseudo-labels, and the target predictor is used to adapt to the data distribution of the target domain during the self-training process, and the threshold is dynamically adjusted to improve the generalization ability of the model.

Benefits of technology

The model's bias towards known targets is reduced, the ability to perceive the unseen targets is improved, the label deviation problem is solved, and the training efficiency and convergence speed are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032111A_ABST
    Figure CN120032111A_ABST
Patent Text Reader

Abstract

The invention relates to an unsupervised domain adaptive open world target detection method, and the method comprises the steps: obtaining a to-be-detected image, and generating a preliminary pseudo tag; determining a source domain and a target domain according to the preliminary pseudo tag, and inputting the source domain and the target domain into a target predictor for target detection; and the target predictor performs self-training according to the labels in the source domain and the target domain so as to adapt to data distribution of the target domain. According to the method, training of the target domain model is guided step by step through generation of the pseudo tag, and the expression of the model on the target domain is improved step by step without depending on the real tag of the target domain by learning the domain invariant feature between the source domain and the target domain, so that the deviation of the model to a known target is reduced, and the perception ability of the model to an unseen target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and more specifically to an unsupervised domain adaptive open-world target detection method. Background Art

[0002] Object detection is a fundamental problem in the field of computer vision, which aims to identify and locate various objects in images. With the development of deep learning technology, object detection has been widely used in many fields such as autonomous driving, intelligent monitoring, and robot navigation. However, most traditional object detection methods are trained and tested in closed environments, that is, they assume that all object categories in the test set are known during the training phase, which has obvious limitations in open scenes in the real world.

[0003] In an open-world environment, an object detection system may encounter unknown objects that have not been seen during the training phase. These unknown objects may have a significant impact on the performance and safety of the system, especially in safety-sensitive applications such as autonomous driving. To address this problem, Joseph et al. proposed the concept of open-world object detection (OWOD), which aims to enable the detection model to recognize and adapt to unknown objects while recognizing known objects.

[0004] Existing OWOD methods try to detect unknown objects through various strategies, such as using detectors trained with known categories to label unknown objects. Although this has improved the performance of object detection in open-world environments to a certain extent, existing methods often rely on the supervised information of known objects to detect unknown objects, which may lead to biased detection of unknown objects by the model, especially for those unknown objects that are not similar or completely unrelated to the features of known objects.

[0005] Secondly, in real open scenarios, after the model detects an unknown object, it will pass it to an annotator (Oracle) to annotate some unknown objects according to its own preferences, and then perform incremental learning based on the new annotation information, so that the model can quickly adapt to and learn new environments. However, during the incremental learning process, the model may forget the previously learned knowledge because when learning new categories, the model needs to adjust its parameters, thereby destroying the previously learned feature representations.

[0006] Therefore, a more effective method is still needed to solve the above problems in open world object detection. Summary of the invention

[0007] In view of this, in order to at least partially solve the above technical problems, the present invention provides an open-world target detection method with unsupervised domain adaptation, aiming to achieve efficient and accurate detection of unknown objects under the background of unsupervised domain adaptation.

[0008] In order to achieve the above object, the present invention adopts the following technical solution:

[0009] An unsupervised domain adaptive open-world object detection method, comprising:

[0010] Get the image to be detected and generate preliminary pseudo labels;

[0011] Construct the source domain based on the true value box and the candidate box with the lowest score in the preliminary pseudo-label, and construct the target domain based on the remaining candidate boxes in the preliminary pseudo-label;

[0012] The source domain and the target domain are input into a target predictor for target detection; the target predictor is self-trained according to labels in the source domain and the target domain to adapt to the data distribution of the target domain.

[0013] Preferably, the self-training step comprises:

[0014] Weakly enhance and strongly enhance the labels in the target domain respectively;

[0015] The target predictor is used to predict the enhanced results and the labels in the source domain, and self-training is performed according to the following loss function based on the prediction results;

[0016] L sum =L SD +λL TD

[0017] In the formula, λ is a hyperparameter, L SD is the cross entropy loss function of the source domain, L TD is the cross entropy loss function of the target domain.

[0018] Preferably, L SD =-[Y SD log(Y SDGT )+(1-Y SD )log(1-Y SDGT )]

[0019] Y SD Represents the prediction result of the label in the source domain, Y SDGT Represents the true value of the label in the source domain.

[0020] Preferably, L TD =-[Y TDs log(Y TDw )+(1-Y TDs )log(1-Y TDw )]

[0021] Y TDs represents the prediction result of the strong enhancement label in the target domain, Y TDw Represents the prediction result of the weakly enhanced label in the target domain.

[0022] Preferably, the prediction result of the weakly enhanced label in the target domain is further predicted based on the threshold, and the expression is:

[0023]

[0024] represents the prediction result of the weakly enhanced label in the target domain, is the threshold value;

[0025] Calculate the cross entropy loss function of the target domain based on the further predicted labels:

[0026]

[0027] To further predict the predicted labels, Represents the prediction result of the strong enhanced label in the target domain.

[0028] Preferably, the threshold is dynamically adjusted according to the evaluation accuracy of each category, and the expression is:

[0029]

[0030] is the flexible threshold at time step t, f t (c) is the estimated accuracy of category c; and

[0031]

[0032] Preferably, f t (c) Mapped to (0, 1) to facilitate adjustment of the fixed threshold; the formula is:

[0033]

[0034] β t (c) represents the scaling factor of category c at time step t.

[0035] Preferably, f t (c) Mapped to (0, 1) to facilitate adjustment of the fixed threshold; the formula is:

[0036]

[0037] β t (c) represents the scaling factor of category c at time step t, and N represents the total number of unlabeled samples.

[0038] Preferably, redundant bounding boxes are filtered by NMS; and potential pseudo-label bounding boxes are identified by IOU filtering; the expression is

[0039]

[0040] Bb m represents the bounding box, m represents the bounding box number, B sam is the set of pseudo-label bounding boxes, gt n For instances of known classes, IOU Td is the IOU threshold.

[0041] Preferably, aspect ratio filtering is performed on the potential pseudo-label bounding boxes;

[0042]

[0043] Ba min with Ba max is the aspect ratio threshold, is the bounding box Bb n The coordinates of .

[0044] The unsupervised domain adaptive open-world target detection method disclosed in the present invention can reduce the bias of the model towards known targets and improve the perception ability of unseen targets, thereby effectively solving the label bias problem, and can significantly improve training efficiency and accelerate convergence speed.

[0045] Compared with the prior art, the present invention uses the basic model to generate the original pseudo-label of the unknown object, which can improve the self-training performance of the target predictor and promote the comprehensive detection of the unknown object;

[0046] At the same time, the threshold is dynamically adjusted according to the learning status of the model, so that the model can use data efficiently throughout the training process. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0048] Figure 1 It is a schematic diagram of the process framework of the unsupervised domain adaptive open-world object detection method of the present invention;

[0049] Figure 2 This is a comparison diagram of the convergence speed of the present invention;

[0050] Figure 3 This is a comparison chart of the qualitative analysis and reasoning speed of the present invention. The yellow box represents the unknown class and the blue box represents the known class. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] In order to address the label bias problem, an embodiment of the present invention proposes an unsupervised domain adaptation detection method for open-world targets. By designing an unsupervised domain adaptation method, instead of the top-k selection strategy, known objects and the lowest-scoring regions are used as source domains, and the remaining unlabeled regions are proposed as target domains. The purpose is to learn a model in a supervised source domain and generalize it to an unsupervised target domain, thereby reducing the model's bias towards known targets and improving its perception of unseen targets.

[0053] Unsupervised domain adaptation methods gradually guide the training of the target domain model by generating pseudo labels. By learning the domain-invariant features between the source domain and the target domain, the performance of the model in the target domain is gradually improved without relying on the true labels of the target domain, thus achieving feature alignment between the source domain and the target domain.

[0054] In one embodiment, the unsupervised domain adaptive open-world object detection method comprises the following steps:

[0055] Get the image to be detected and generate preliminary pseudo labels;

[0056] Construct the source domain based on the true value box and the candidate box with the lowest score in the preliminary pseudo-label, and construct the target domain based on the remaining candidate boxes in the preliminary pseudo-label;

[0057] The source domain and the target domain are input into a target predictor for target detection; the target predictor is self-trained according to labels in the source domain and the target domain to adapt to the data distribution of the target domain.

[0058] Reference Figure 1 , Figure 1 This is a schematic diagram of the process framework of the unsupervised domain adaptive open-world object detection method of this application;

[0059] In one embodiment, in open world object detection, there are limited unknown object identifiers, and a large number of unknown object labels will improve the model's recall ability. However, since it is difficult to distinguish unknown objects from the background, it is difficult to label unknown objects. Therefore, this application proposes a method for generating pseudo labels for unknown objects based on a basic model. It includes:

[0060] Firstly, the basic model is used to generate preliminary pseudo labels of unknown objects according to the images to be detected for auxiliary supervision, thereby improving the self-training performance of the dynamic learning based unsupervised domain adaptation strategy (DLDUA), which should promote the comprehensive detection of unknown objects.

[0061] With the rapid development of the basic model, it has shown significant superiority. Therefore, this embodiment uses SAM to generate preliminary pseudo labels for unknown objects.

[0062] like Figure 1 In (a), this application uses the ability of SAM (Segment Anything Mode) to segment all objects and fully segment a single image M to generate a class-agnostic mask:

[0063]

[0064] Among them, for the current mask, m sam and m ms Represents the mask and number of pixels.

[0065] Then, we calculate the maximum x and y coordinates for each mask to generate a set of bounding boxes:

[0066]

[0067] Since the generated bounding box contains GT (ground truth) information and noise, GT refers to the real label or data used by the model in training and testing; therefore, this embodiment uses non-maximum suppression (NMS) and IOU (Intersection over Union) to filter out potential unknown objects, including:

[0068] Use NMS to filter duplicate bounding boxes.

[0069] Calculate the IOU between the bounding box generated by SAM and the annotated object. If the IOU is less than the set threshold IOU Td , it is considered an unknown object:

[0070]

[0071] Among them, Bb m represents the bounding box, m represents the bounding box number, B sam is the set of pseudo-label bounding boxes, gt n For instances of known classes, IOU Td is the IOU threshold.

[0072] Furthermore, SAM often generates new masks for the internal pixels of the segmented objects. If a mask is too large or too small, it may be caused by noise. Therefore, this application further filters by the aspect ratio of the bounding box. If the aspect ratio of the border of an object is between Ba min with Ba max , then we know it as the original pseudo label of the denoised unknown object:

[0073]

[0074] Ba min with Ba max is the aspect ratio threshold, is the bounding box Bb n The coordinates of .

[0075] The generated rough pseudo-labels of unknown objects are then used for auxiliary supervision to improve the predictor self-training performance, which should facilitate comprehensive detection of unknown objects.

[0076] In one embodiment, a source domain is constructed based on the true value box and the candidate box with the lowest score in the preliminary pseudo-label, and a target domain is constructed based on the remaining candidate boxes in the preliminary pseudo-label;

[0077] In this embodiment, the open world object detection framework is as follows Figure 1 As shown in the upper part, given an input image, after passing through the basic model of known classes, the output includes known category detection results (green boxes) and multiple unknown category candidate regions with object scores (yellow boxes). Furthermore, the unmatched candidate regions are sorted, the true value box label is marked as 1, and the candidate region with the lowest score is marked as 0. The two together constitute the source domain; the remaining unmatched candidate regions are used as the target domain.

[0078] In one embodiment, if Figure 1 In (c), the source domain and the target domain are input into the target predictor together, and the target predictor is self-trained to finally output the detection result of whether the candidate region is an unknown object.

[0079] In a specific embodiment, an unsupervised domain adaptation method is used to train the target predictor, thereby obtaining a foreground background predictor that can generalize in an unlabeled target domain. The training step includes:

[0080] First, the labels in the target domain are weakly enhanced to obtain TD w ; Then it is strongly enhanced to get TD s ;

[0081] Secondly, the target predictor is used to predict the target of the enhanced result and the label in the source domain.

[0082] Reference Figure 1 In (b) of this embodiment, the target predictor is a classification network composed of a multi-scale convolutional layer and a fully connected layer, which is used to determine whether the input is a foreground target or a background region, denoted by φ; the corresponding prediction result is:

[0083] Y SD = φ(SD), Y TDW = φ(TD w )), Y TDs = φ(TD s )

[0084] Y SD 、Y TDW 、Y TDs are the outputs after the model φ processes the input data, that is, logits. These output logits are then used to calculate the loss function to guide the model training; among them, the total loss function of unsupervised domain adaptation is expressed as;

[0085] L sum = L SD + λL TD

[0086] In the formula, λ is a hyperparameter used to adjust the relative importance between the two losses to optimize the generalization ability of the model in the target domain. L SD is the cross-entropy loss function of the source domain, and L TD is the cross-entropy loss function of the target domain.

[0087] According to consistency regularization, during the model training process, L SD is used to represent the cross-entropy loss function of the source domain and is used to calculate the error between the model's prediction Y SD and the true label Y SDGT :

[0088] L SD = -[Y SD log(Y SDGT ) + (1 - Y SD ) log(1 - Y SDGT )]

[0089] Y SD represents the prediction result of the label in the source domain, and Y SDGT represents the true value of the label in the source domain. Y SD log(Y SDGT ) represents the loss of the foreground label, and (1 - Y SD ) log(1 - Y SDGT ) represents the loss of the background.

[0090] According to consistency regularization, during the model training process, L TDThe cross entropy loss function used to represent the pseudo-label prediction in the target domain measures the pseudo-label Y predicted by the model. TDw and the enhanced pseudo-label Y TDs Consistency between:

[0091] L TD =-[Y TDs log(Y TDw )+(1-Y TDs )log(1-Y TDw )]

[0092] Y TDs represents the prediction result of the strong enhancement label in the target domain, Y TDw represents the prediction result of the weakly enhanced label in the target domain, Y TDs log(Y TDw ) represents the loss of the foreground label, (1-Y TDs )log(1-Y TDw ) represents the background loss.

[0093] The above self-training methods usually rely on a fixed threshold to generate artificial pseudo-labels, and use unlabeled data with prediction confidence higher than the fixed threshold for unsupervised training. Although this strategy can ensure that only high-quality unlabeled data is used for model training, it ignores a large amount of other unlabeled data, especially in the early stages of model training, when only a small amount of unlabeled data is higher than the fixed threshold. In addition, the learning difficulty differences between different categories are not considered.

[0094] Therefore, in order to reduce the possibility of misclassification and avoid some samples with large uncertainty from being mistakenly labeled as foreground or background, the present application provides an unsupervised domain adaptation strategy (DLDUA) based on dynamic learning, which dynamically adjusts the threshold according to the learning state of the model, especially in the early stages of model training, greatly increasing the proportion of effective unlabeled samples, so that the model can use data efficiently throughout the entire training process.

[0095] In a specific embodiment, according to the paradigm of pseudo-label generation, the prediction results of the weakly enhanced labels in the target domain are further predicted based on the threshold, and the expression is:

[0096]

[0097] represents the prediction result of the weakly enhanced label in the target domain, is the threshold value;

[0098] In this embodiment, the pseudo label The predicted results are calculated The maximum value of the softmax output is at the threshold Generated by comparison, if the maximum value of softmax exceeds the threshold The generated pseudo label represents the foreground target, otherwise it is the background area.

[0099] Then, the cross entropy loss function of the target domain is calculated based on the further predicted labels:

[0100]

[0101] To further predict the predicted labels, Represents the prediction result of the strong enhanced label in the target domain.

[0102] This application dynamically adjusts the threshold according to the learning state of the model, so that the model can use data efficiently throughout the entire training process.

[0103] In this embodiment, for the pseudo-label generation process, pseudo-labels are generated only based on fixed high-confidence unlabeled data to calculate the unsupervised loss. Although this strategy can ensure that only high-confidence unlabeled data is used for model training, it also ignores a large amount of unlabeled data, especially in the early stage of model training, when only a small amount of unlabeled data is above the threshold.

[0104] Therefore, according to the pseudo label (CPL) strategy, the learning state of each category is considered. The core idea is to let the model learn tasks sequentially from simple to complex, which can improve the training efficiency and generalization of the model. However, it is difficult to determine the threshold according to the learning state, so the present invention dynamically adjusts the threshold according to the evaluation accuracy of each category.

[0105] In an exemplary embodiment, the threshold value is dynamically adjusted as follows:

[0106]

[0107] Where c∈{1,...,B,U k} represents the known and unknown categories in the current task, represents the flexible threshold at time step t, f t (c) is the corresponding evaluation accuracy, such that the lower the accuracy, the worse the learning state of the model, resulting in a decrease in the threshold, thereby encouraging more unlabeled data to participate in model training.

[0108] According to the CPL strategy, it is assumed that when the threshold is high, the learning effect of a category can be reflected by the number of samples predicted to fall into the category and above the threshold. In other words, the category with a small number of samples whose prediction confidence reaches the threshold is considered to have greater learning difficulty or poor learning status; that is,

[0109]

[0110] Where N is the total number of unlabeled data. The learning effect of category c is determined by the number of unlabeled samples that meet two conditions: Indicates that the model's prediction confidence for these samples exceeds the threshold. Indicates that the predicted category of these samples is c.

[0111] The dynamic learning-driven unsupervised domain adaptation strategy proposed in this application can reduce the model's bias towards known targets and improve the recall rate of unknown targets.

[0112] To further optimize the above technical solution, in one embodiment, by t (c) Perform normalization so that its mapping range is between (0, 1) to adjust the fixed threshold:

[0113]

[0114] f t represents the learning effect, β t (c) represents the scaling factor of category c at time step t, which is calculated by subtracting the confidence f of category c from t (c) Calculated by dividing by the maximum confidence among all categories.

[0115] Right now

[0116] This formula defines the adaptive confidence threshold for category c It is the scaling factor β t (c) and preset constant threshold The product of .

[0117] Furthermore, due to the influence of model parameterization in the early stage of training, blind prediction is prone to occur, resulting in label bias. That is, the model tends to classify most unlabeled samples into a specific category, making the early learning state unreliable. Therefore, a mitigation process is introduced:

[0118]

[0119] Among them, f t (c) represents the learning effect of category c at time step t, N represents the total number of unlabeled samples, represents the value of the maximum learning effect in all categories at time step t, represents the sum of the learning effects of all categories at time step t, It can be regarded as the amount of unlabeled data that is not used. To avoid excessive bias in the early stages of the model.

[0120] In this formula, the denominator acts as a safeguard in the early stages of training, reducing certain deviations caused by unreliable learning states.

[0121] The entire process described above in this application does not introduce any additional parameters (such as hyperparameters or trainable parameters) or additional calculations, and therefore does not increase the training overhead of the model. On the contrary, due to the setting of the dynamic threshold, the model training is smoother and the convergence speed of the model is accelerated.

[0122] In one embodiment, the detection process includes:

[0123] 1. Initial training: First, train the model on the source domain (labeled data) to obtain the initial model parameters.

[0124] 2. Generate pseudo labels: Apply the trained model to the target domain (unlabeled data) and generate pseudo labels based on the model’s predictions; pseudo labels are the model’s predicted categories for the target domain data.

[0125] 3. Self-training: Use the target domain data with pseudo labels to further train the model. The model uses pseudo labels to adjust its parameters through step-by-step iterations, so that it gradually adapts to the data distribution of the target domain.

[0126] 4. Repeated iteration: The self-training process can be repeated to gradually optimize the model so that it has better generalization ability in the target domain.

[0127] Furthermore, the present invention has conducted comprehensive experiments on the open-world object detection benchmark, which has verified the superiority of the model of the present invention compared with the existing SOTA (state-of-the-art) method.

[0128] For the fairness of the experiment and consistency with previous work, this application uses VOC and COCO to form a mixed dataset. In order to simulate the incremental learning process, the dataset is divided into four tasks. The training images and instances contained in each task are as follows:

[0129] T1 (16551, 47223), T2 (45520, 113741), T3 (39402, 114452), T4 (40260, 138996)

[0130] Among them, task one includes all 20 categories of voc, and other tasks add 20 new categories on this basis. The test set contains 10,246 images and 61,707 test instances in all categories. During the evaluation process, all categories seen earlier (including new categories) are considered known categories, and the remaining categories are considered unknown categories.

[0131] For the detection of known categories, the mean average precision (mAP) is used as the evaluation indicator. During the incremental learning process (T2-T4), mAP is further subdivided into Previously known mAP, Current known mAP and Both mAP; Previously known mAP is used to evaluate the impact of newly added categories on existing knowledge in incremental learning tasks, Current known mAP indicates the learning effect of the model on new categories at the current stage, and Both mAP evaluates the overall detection performance of the model on all known categories.

[0132] For the detection of unknown categories, U-Recall is used as the main evaluation metric, which represents the ability to detect unknown objects. In addition, in order to evaluate the degree of confusion between known and unknown categories, WI and A-OSE are used for evaluation.

[0133] For the prediction model, the experiment is based on the Faster-RCNN architecture. In the source domain, the number of true labels and credible background ratios are set equal, and the same number of proposals are extracted from the mismatched proposals as the target domain. The initial threshold in UDA is set to 0.95, and λ1=1, λ2=0.3 in the loss function; at the same time, DINO pre-trained restnet-50 is used as the backbone network. For model training, the initial learning rate is 0.005, the batch_size is 8, and 4 NVIDIA RTX 3090GPUs are used. The code is implemented based on detectron2.

[0134] For the generation of the initial pseudo-label, set the IOU Td The value is 0.6, which is used to filter the true value box, and the aspect ratio is used to filter the noise box. The bounding box that does not meet the aspect ratio between 0.3 and 4 is considered to be a noise box, and finally the original pseudo label of the unknown object is obtained.

[0135] In order to reflect the advantages of this application, the latest representative related works of OWOD are compared respectively. Compared with the original ORE, these methods have achieved new SOTA. In these works, the energy model (EUBI) in ORE is removed, so it will cause data leakage problems; for the fairness of the experiment, the method SGROD that also uses the basic model SAM is further compared, but the method SKDF that uses the basic model clip on OWOD is not compared, because it is trained on large-scale data and reasoned on the OWOD dataset, which has serious data leakage problems. The method in this application only uses SAM to generate preliminary pseudo-labels for unknown objects, and does not know the known and unknown object categories in advance, so there is no data leakage problem. The experimental data refers to Table 1;

[0136] Table 1

[0137]

[0138] As can be seen from Table 1, in terms of identifying known and potential unknown objects, in terms of unknown category recall in Task 1, the method of the present invention is ahead of the method ORTH which does not use a large model by 15.2 U-Recall, and ahead of the method SGROD which uses a large model by 5.5 U-Recall. In terms of known class detection methods, it is ahead of the method SGROD which uses a large model by 0.4 mAP, and ahead of ORTH by 0.2 mAP. In terms of incremental learning, in T2-T4, this application has achieved comprehensive leadership, such as leading SGROD by 18.7% (T2: 38.7 vs 32.6), leading ORTH by 47.1% (T2: 38.7 vs 26.3), leading SGROD by 8.9% (T3: 35.6 vs 32.7), leading ORTH by 22.3% (T3: 35.6 vs 29.1), and leading SGROD by 6.3 (T2: 38.6 vs 32.3) for known category mAP, leading SGROD by 9.8 (T3: 32.2 vs 22.4) and leading SGROD by 8.2 (T4: 26.7 vs 18.5). It can be seen from the experimental results that the dynamic learning method proposed in the present invention can improve the detection performance of known and unknown classes, and alleviate the risk of known detection performance degradation during incremental learning.

[0139] Furthermore, Table 2 compares the confusion between known and unknown categories.

[0140] Table 2

[0141]

[0142] Please note that since all 80 classes in Task 4 are known, experiments are only conducted on Tasks 1 to 3. From the experimental results, it can be seen that in addition to having a huge advantage in unknown recall, the WI and A-OSE are still at the state-of-the-art performance. It can be seen that the present invention can not only recall more unknown objects, but also distinguish known and unknown categories well, thanks to the use of the basic model and dynamic learning strategy.

[0143] Table 3

[0144]

[0145] Referring to Table 3, in Task 2, the foreground / background ratios in the source domain are set to 1:1, 1:2, 1:5, and 1:10. From the experimental results, it can be seen that the experiment with a foreground and background ratio of 1:1 in the source domain has the best effect. Because this ratio helps the model better capture and distinguish the features of foreground objects and the background in domain adaptation. A balanced foreground and background ratio avoids the problem of class imbalance, reduces the overfitting of the model to the background, and at the same time enhances the recognition ability of foreground objects. This way can help the model better adapt to the detection task of unknown objects in the target domain.

[0146] Furthermore, referring to Table 4,

[0147] Table 4

[0148]

[0149] In Task 2, different values of the parameter λ in the loss function are set. The parameter λ1 is used to adjust the relative importance between the training losses of the source domain and the target domain. A balanced ratio helps the model better learn the features of both. An unbalanced ratio may cause the model to deviate from the features of certain classes, thus affecting the detection of known and unknown classes. If the parameter λ2 is set too large, the model may pay more attention to the target domain, resulting in a decline in performance on the source domain. By appropriately controlling λ2, this phenomenon of "over domain adaptation" can be avoided. From the above and the experimental results, it can be seen that λ1 = 1 and λ2 = 0.4 is a reasonable balance point.

[0150] The adjustment of the flexible threshold is achieved through a normalization function. Since in the early stage, the training of the model is not stable enough, large fluctuations will occur in the increase and decrease of β t (c), but it will tend to be stable in the middle and late stages of training. Therefore, the choice of the normalization function is crucial. As shown in Table 5,

[0151] Table 5

[0152]

[0153] By comparing four normalization functions: (1) Concave function: M(x) = ln(x + 1) / ln2, (2) Convex function: M(x) = x / (2 - x), (3) Linear function: M(x) = x, and (4) Logistic function: M(x) = 1 / (1 + e^(-x)), the results show that the linear and concave functions perform poorly, while the convex and logistic functions perform better. The convex and logistic functions have advantages in terms of early training stability and mid - to - late - stage fitting effect, probably because their non - linear smoothing characteristics are more suitable for the dynamic characteristics of pseudo - label learning.

[0154] Regarding the generation of position object region proposals, as shown in Table 6,

[0155] Table 6

[0156]

[0157] This application evaluates different unknown object region proposal generation methods, including Freesolo for unsupervised instance segmentation, DETReg based on self-supervised pre-training, EdgeBoxes based on contour segmentation, Geodesic Object Proposals (GOP) based on low-level visual information, and SAM method. These methods are compared with a baseline that does not use any unknown object proposal generation. From the experimental results, it can be seen that in the unknown recall method, all region proposal generation methods produce better results than the baseline; although the traditional method GOP can only generate rough unknown bounding boxes, the dynamic learning strategy of the present invention can still improve known and unknown detections from noisy environments.

[0158] For dynamic thresholds, the present invention compares the performance of dynamic thresholds and fixed thresholds in terms of known category detection accuracy and unknown category recall on multiple tasks. The experimental results are shown in Table 7.

[0159] Table 7

[0160]

[0161] It can be seen that the dynamic threshold method can improve the accuracy of known category detection and improve the recall rate of unknown categories, especially in tasks with less labeled data. This is due to the fact that the dynamic learning method can dynamically adjust the threshold according to the learning effect of the model without causing the problem of threshold truncation, and further solves the label bias problem.

[0162] Another advantage of dynamic learning is the convergence speed, e.g. Figure 2 As shown in the figure, by comparing the dynamic threshold and the fixed threshold in terms of loss and detection accuracy. The dynamic learning method makes the loss drop faster and smoother, showing an excellent convergence speed; in contrast, the loss of the fixed threshold learning method fluctuates greatly, probably because the fixed threshold makes most of the unlabeled data of certain categories fail to pass, and the present application can make more unlabeled data pass, so that the gradient reaches the global optimum.

[0163] Figure 3The results of qualitative analysis and reasoning speed of the model of the present invention are shown. Compared with the latest work ORTH that does not use a large model and the SGROD method that uses a large model, in terms of qualitative analysis, the present application can not only detect a considerable number of known objects, but also recall more unknown objects, such as the first row from top to bottom: the present invention can detect the white cabinets missed by SGROD and ORTH; the second row can detect the lamps, stools, black cabinets, etc. that SGROD and ORTH missed, so it can be concluded that the method provided by the present application performs better in the OWOD scene. In terms of reasoning speed, using a single NVIDIA RTX3090 GPUs for reasoning, the method of the present invention is ahead of SGROD by 10.06 FPS (34.56 vs 24.51) and ORTH by 2.72 FPS (34.56 vs 31.84). This is mainly because compared with SGROD, the present application uses a pure convolutional neural network to build, while SGROD uses a dense attention mechanism. Therefore, the reasoning of the present invention is faster. Compared with ORTH, although both are based on the faster-rcnn architecture, ORTH uses more parameters and introduces the constraints of orthogonal matrices, which increases the computational complexity and leads to slower reasoning. It can be concluded that the present invention not only has faster reasoning, but also has the potential for real-time applications.

[0164] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0165] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An unsupervised domain adaptive open-world object detection method, characterized in that: Get the image to be detected and generate preliminary pseudo labels; Construct the source domain based on the true value box and the candidate box with the lowest score in the preliminary pseudo-label, and construct the target domain based on the remaining candidate boxes in the preliminary pseudo-label; Input the source domain and target domain into the target predictor for target detection; The target predictor is self-trained based on the labels in the source and target domains to adapt to the data distribution of the target domain.

2. The open world object detection method according to claim 1, characterized in that: The self-training steps include: Weakly enhance and strongly enhance the labels in the target domain respectively; The target predictor is used to predict the enhanced results and the labels in the source domain, and self-training is performed according to the following loss function based on the prediction results; THE sum =L SD +λL TD In the formula, λ is a hyperparameter, L SD is the cross entropy loss function of the source domain, L TD is the cross entropy loss function of the target domain.

3. The open world object detection method according to claim 2, characterized in that: L SD =-[And SD log(Y SDGT )+(1-Y SD )log(1-Y SDGT )] Where Y SD Represents the prediction result of the label in the source domain, Y SDGT Represents the true value of the label in the source domain.

4. The open world object detection method according to claim 2, characterized in that: L TD =-[And TDs log(Y TDw )+(1-Y TDs )log(1-Y TDw )] Where Y TDs represents the prediction result of the strong enhanced label in the target domain, Y TDw Represents the prediction result of the weakly enhanced label in the target domain.

5. The open world object detection method according to claim 2, characterized in that: The prediction results of the weakly enhanced labels in the target domain are further predicted based on the threshold, and the expression is: In the formula, represents the prediction result of the weakly enhanced label in the target domain, is the threshold value; Calculate the cross entropy loss function of the target domain based on the further predicted labels: In the formula, To further predict the predicted labels, Represents the prediction result of the strong enhanced label in the target domain.

6. The open world object detection method according to claim 5, characterized in that: The threshold is dynamically adjusted according to the evaluation accuracy of each category, and the expression is: In the formula, is the flexible threshold at time step t, f t (c) is the estimated accuracy of category c; and 7. The open world object detection method according to claim 6, characterized in that: f t (c) Mapped to (0, 1) to facilitate adjustment of the fixed threshold; the formula is: In the formula, f t represents the learning effect, β t (c) represents the scaling factor of category c at time step t.

8. The open world object detection method according to claim 6, characterized in that: f t (c) Mapped to (0, 1) to facilitate adjustment of the fixed threshold; the formula is: In the formula, β t (c) represents the scaling factor of category c at time step t, and N represents the total number of unlabeled samples.

9. The open world object detection method according to claim 1, characterized in that: Redundant bounding boxes are filtered through NMS; and potential pseudo-label bounding boxes are identified by filtering through IOU; the expression is In the formula, Bb m represents the bounding box, m represents the bounding box number, B sam is the set of pseudo-label bounding boxes, gt n For instances of known classes, IOU Td is the IOU threshold.

10. The open world object detection method according to claim 9, characterized in that: Perform aspect ratio filtering on potential pseudo-label bounding boxes; In the formula, Ba min with Ba max is the aspect ratio threshold, is the bounding box Bb n The coordinates of .

Citation Information

Patent Citations

  • Open world target detection method based on visual large model enhancement

    CN118097289A

  • Unknown classifiable open world target detection method

    CN118747795A

  • Method and device for domain adaptation training of neural network

    JP2024125219A