Open world target detection method based on progressive feature fusion and uncertainty estimation

By adopting techniques of progressive feature fusion and uncertain estimation in the open-world object detection model, the confusion, low recall and forgetting problems of the model when facing unknown categories are solved, achieving higher detection accuracy and recall.

CN120014350AActive Publication Date: 2025-05-16CHINA UNIV OF MINING & TECH

Patent Information

Application Number
CN202510100928.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Open-world object detection models have problems with known categories and unknown categories when facing unknown categories, low recall rates for unknown categories, and catastrophic forgetting.

Method used

Using a method based on progressive feature fusion and uncertainty estimation, a multi-scale feature map was extracted and weighted fusion was performed through the improved Fast R-CNN network and the ResNet-50 backbone network, uncertainty estimation was performed, and the category classification score was adjusted through the improved Fast R-CNN network and the ResNet-50 backbone network.

Benefits of technology

It improves the accuracy of detection of known categories, enhances the recall rate of unknown categories, reduces the model's dependence on manual annotation, and improves the robustness and applicability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014350A_ABST
    Figure CN120014350A_ABST
Patent Text Reader

Abstract

The invention discloses an open world target detection method based on progressive feature fusion and uncertainty estimation, and the method comprises the steps: detecting a natural image through employing an open world target detection model, and determining the coordinates and confidence of a known class and an unknown class in the image; wherein the open world target detection model comprises the following steps: acquiring an open world target detection data set, and carrying out data preprocessing; the method comprises the following steps: constructing an open world target detection model by using Fast R-CNN, and extracting initial multi-scale features by ResNet-50; gradually fusing the initial multi-scale features to obtain fused multi-scale features; carrying out attention calculation on the fused multi-scale features to obtain fixed features; predicting classification scores of the fixed features; calculating the classification score to obtain a variance for uncertainty estimation; judging whether the class belongs to a known class or an unknown class based on the uncertainty and the classification score; performing backdoor adjustment based on uncertainty to obtain a final classification score; and training the open world target detection model by using the open world target detection data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an open world target detection method based on progressive feature fusion and uncertainty estimation, and belongs to the technical field of image processing. Background Art

[0002] With the rapid development of deep learning technology, deep learning has been applied to more and more fields. In the field of natural image processing, there is a large demand for natural image processing, so computer vision technology based on deep learning has been widely used in various tasks in this field. The proposal of open-world object detection is to solve the limitations of traditional object detection methods when facing new environments and new objects. In traditional object detection, all categories that need to be detected are provided with labels during the training phase, while open-world object detection relaxes this premise, allowing the model to encounter natural images of unknown categories during the testing phase, and be able to detect and learn these unknown categories.

[0003] Open-world object detection is a fundamental task in the field of computer vision. The main goal of this task is to scan natural images for objects that may need to be identified, find all known categories and possible unknown categories, and mark the location and category of each object. The model framework of open-world object detection needs to recognize objects of known categories and gradually adapt to unknown categories, using techniques such as incremental learning and self-supervised learning to improve detection capabilities in an ever-expanding dynamic environment.

[0004] In the traditional closed world, the existing object detection frameworks are relatively complete. However, there are still some challenges in applying these frameworks to the open world. The reason is that without effective supervision of unknown categories during the training process, the model will mistakenly identify known categories as unknown categories; secondly, when new categories are constantly introduced, the model is prone to forget previously learned knowledge; at the same time, the open world environment requires the model to be able to autonomously discover and learn new categories, thereby reducing dependence on a large amount of manual annotation. The challenges faced by open world object detection are as follows:

[0005] (1) Confusion between known and unknown categories: Due to the limitations of training data, feature similarity, and the lack of an identification mechanism for “unknown” categories, the model lacks accuracy when identifying new categories. This not only affects the detection effect, but also reduces the robustness and decision reliability of the model in real scenarios.

[0006] (2) Low recall rate of unknown categories: Since the model is usually trained only on known categories, its recognition and detection capabilities are limited when encountering unknown categories, resulting in the inability to effectively identify all new category objects. This low recall rate problem stems from the model’s lack of sensitivity to the features of unknown categories, especially when there are obvious differences in feature distribution from known categories. The model has difficulty identifying these objects as independent new categories, thus affecting the comprehensiveness and applicability of detection.

[0007] (3) Catastrophic forgetting: In an open world environment, the model needs to constantly adapt to newly added categories, but the incremental learning process lacks repeated training of the original categories, which causes the model's recognition accuracy to drop significantly when new categories are added. This phenomenon causes the model to lose its ability to recognize the original categories while expanding new categories, affecting the continuity and accuracy of detection. Summary of the invention

[0008] The technical problem to be solved by the present invention is to overcome the defects of the prior art, provide an open-world target detection method based on progressive feature fusion and uncertainty estimation, and propose corresponding solutions to the problems of confusion between known classes and unknown classes and low recall rate of unknown classes.

[0009] Preferably, the present invention provides an open world object detection method based on progressive feature fusion and uncertainty estimation, characterized by comprising:

[0010] Use the trained open-world object detection model to detect natural images of natural scenes, and determine the classification results of known category features, the location information of known category features, the target scores of known category features, the classification results of unknown category features, the location information of unknown category features, and the target scores of unknown category features;

[0011] Among them, the training of the open world target detection model includes:

[0012] Step 1, obtaining an open world object detection dataset and preprocessing the data in the open world object detection dataset;

[0013] Step 2: Use the improved Fast R-CNN network to build an open-world object detection model. The open-world object detection model uses ResNet-50 as the backbone network to extract the initial multi-scale feature map; perform weighted progressive fusion on the initial multi-scale feature map to obtain a fused multi-scale feature map; perform cross-attention calculation on the fused multi-scale feature map to obtain a fixed-size feature map; perform category classification score, bounding box regression score and target score on the fixed-size feature map;

[0014] Step 3: Make an uncertainty estimate of the category classification score, use Monte Carlo sampling to calculate the variance of the category classification score as the confidence of the classification result, and quantify the uncertainty based on the confidence; judge whether the category corresponding to the category classification score belongs to a known category or an unknown category based on the uncertainty; and make a backdoor adjustment based on the uncertainty to obtain the final category classification score;

[0015] Step 4: Use the open-world object detection dataset to train the open-world object detection model.

[0016] Preferably, step 1, obtaining an open world object detection dataset and preprocessing the data in the open world object detection dataset, includes:

[0017] Step (1.1), obtaining an open-world object detection dataset, the open-world object detection dataset contains natural images with labeled object instances, and the open-world object detection dataset has multiple known categories and several unknown categories;

[0018] Step (1.2), cropping the natural image in the open-world object detection dataset to obtain an image of a preset size; modifying the instance information in the open-world object detection dataset to adapt to the image of the preset size;

[0019] Step (1.3) is to label the graph of preset size, redivide it into several tasks, and perform incremental learning on the tasks.

[0020] Preferably, step 2 comprises:

[0021] Step (2.1), define the feature maps of different scales of the backbone network Resnet50 as {C1, C2, C3, C4, C5}, where C1 corresponds to the features of the initial convolutional layer and pooling layer, and C2, C3, C4, C5 correspond to the features of subsequent layers;

[0022] Step (2.2), substitute C1 into a convolutional layer, batch normalization layer and ReLU activation function for processing to obtain C2;

[0023] Substitute C2 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C3;

[0024] Substitute C3 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C4;

[0025] Substitute C4 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C5;

[0026] Step (2.3), use global maximum pooling for C1, reduce the width and height of C1 to 1 / 4 of the original, apply convolution with a step size of 1 and 1x1 to C2, and obtain the fused feature map P2;

[0027] Step (2.4), apply 1x1 convolution and bilinear interpolation to upsample C2 to align the spatial size of C2 with C3 to obtain C2′; apply 2x2 convolution with a step size of 2 to C3 to align the feature dimension of C3 with C2′;

[0028] Step (2.5), calculate the weight of C2′ and the weight of C3 and perform weighted fusion to obtain the fused feature map P3;

[0029] Step (2.6), repeat the above steps (2.4) and (2.5), weightedly fuse the features of C3 and C4, and the features of C4 and C5; when weighted fusing the features of C3 and C4, apply convolution with a step size of 4 and 4x4 to obtain the fused feature map P4; weighted fusing the features of C4 and C5, apply convolution with a step size of 8 and 8x8 to obtain the fused feature map P5, and obtain a set of fused multi-scale feature maps {P2, P3, P4, P5};

[0030] In step (2.7), the fused multi-scale feature map is cross-attention calculated to obtain a fixed-size feature map; the fixed-size feature map is scored for category classification, bounding box regression, and targetness.

[0031] Preferably, step (2.6) comprises:

[0032] For each pair of data in the feature map {(C1, C2), (C2, C3), (C3, C4), (C4, C5)}, perform the following steps:

[0033] Step (2.41), apply a convolution layer, a batch normalization layer and a ReLU activation function to a pair of feature maps respectively to obtain the first feature and the second feature;

[0034] Step (2.42), concatenate the first feature and the second feature in the channel dimension to obtain a third feature map;

[0035] Step (2.43), apply 1x1 convolution to the third feature map, and calculate the weights of feature maps at all levels;

[0036] Step (2.44), use the softmax function to normalize the weights of feature maps at all levels to obtain the fusion weight of each feature map;

[0037] Step (2.45), according to the fusion weight of each feature map, the feature maps of different scales are weightedly fused to obtain a fused feature map;

[0038] Step (2.46), apply a convolution layer, batch normalization layer and ReLU activation function to the fused feature map to obtain the final weighted fused feature map;

[0039] In step (2.47), all the final weighted fusion feature maps are combined to obtain the final weighted multi-scale fusion feature map {P2, P3, P4, P5}.

[0040] Preferably, in step 3, an uncertainty estimation is performed on the category classification score, and the variance of the category classification score is calculated using Monte Carlo sampling as the confidence of the classification result, and the uncertainty is quantified according to the confidence; based on the uncertainty, it is determined whether the category corresponding to the category classification score belongs to a known category or an unknown category; and a backdoor adjustment is performed based on the uncertainty to obtain the final category classification score, including:

[0041] Step (3.1), fusion feature map F for the i-th sampling i Perform N forward propagations for each pixel in, i = 1, 2, ..., N;

[0042] Step (3.2), the fusion feature map F obtained by the i-th sampling i Use classifier C to classify and get the predicted probability distribution P for each category i (c);

[0043] Step (3.3), calculate the mean μ(c) and variance σ of the predicted probability of each category c 2 (c), obtain the uncertainty; where,

[0044]

[0045] Step (3.4), based on the uncertainty, evaluate the classification score of each natural image to determine whether there is an unknown category; if the classification score of the natural image exceeds a first preset unknown classification threshold, it is determined that there is an unknown category;

[0046] In step (3.5), the calculated uncertainty is used as the uncertainty measure for backdoor adjustment.

[0047] Preferably, step (3) comprises:

[0048] Step (3.51), perform N forward propagations on the category classification scores during the inference process of the open world object detection model;

[0049] Step (3.52), for each category c, calculate the average probability value u(c) of all N Monte Carlo samplings;

[0050] Step (3.53), for each category c, calculate the variance σ of the predicted probability of all N Monte Carlo samples 2 (c);

[0051] Step (3.54), using u(c) and σ 2 (c) evaluating the prediction uncertainty of each natural image in the open world object detection dataset for a specific category, and if the classification score of the natural image exceeds a second preset unknown classification threshold, determining that an unknown classification exists;

[0052] Step (3.55), calculate the variance σ 2 (c) As a confidence uncertainty measure for each natural image in the open-world object detection dataset to dynamically adjust the classification score and of each natural image.

[0053] Preferably, step (3.5) comprises:

[0054] (3.51) If the variance σ 2 (c) if it exceeds a preset threshold τ, the corresponding natural image is judged to have high uncertainty;

[0055] (3.52) The variance value σ of each high uncertainty natural image 2 (c) As an adjustment factor, it is added to the classification score of natural images;

[0056] (3.53) Calculate the classification score S′ of the natural image of each high uncertainty sample:

[0057] S′=S+β·σ 2 (c),

[0058] Among them, S is the original classification score, and β is a preset coefficient;

[0059] (3.54) If the backdoor adjusted score S′ of each category c exceeds the adjusted third classification threshold, the corresponding natural image is judged to belong to category c.

[0060] Preferably, step (2.7) performs cross-attention calculation on the fused multi-scale feature map to obtain a fixed-size feature map, including:

[0061] (2.71), receiving the input feature map f, the dimension of the feature map f is (b,n,c), b is the batch size, n is the number of detection boxes, and c is the feature dimension;

[0062] (2.72), the feature dimension dim=64 and the input resolution of the feature map f is 1024;

[0063] (2.73), use the nn.Conv1d(dim, dim, 3, padding = 1, groups = dim) convolution layer to perform depth-separable convolution on the feature map f to obtain the processed feature map; add the processed feature map to the original input feature map to form a residual connection;

[0064] (2.74), normalize the input feature map, then use the linear layer to transform the feature map and apply the activation function, repeat this step twice;

[0065] (2.75), calculate the cross attention based on the input feature map input obtained in step (2.74) to obtain the weighted feature representation;

[0066] (2.75), transform the weighted feature representation through a linear layer, add the weighted feature representation and the residual connection to obtain the feature map y;

[0067] (2.76), use the nn.Conv1d(dim,dim,3,padding=1,groups=dim) convolution layer to perform depth-wise separable convolution on the feature map y;

[0068] (2.77), the input feature map output by step (2.76) is normalized and further processed through the MLP layer to generate an output feature map of fixed size.

[0069] Preferably, step (2.75) comprises:

[0070] Expand each dimension of the obtained input feature map x into (b,n,num_heads,head_dim), where b is the batch size, n is the number of detection boxes, num_heads is the number of attention heads, and head_dim is the dimension of the attention head;

[0071] Perform self-attention calculation on x:

[0072] Q=φ(xW Q ), (1),

[0073] K=φ(xW K ), (2),

[0074] V=φ(xW V ), (3),

[0075]

[0076] Attn normed =Norm(Attn)+x, (5),

[0077] output=Norm(MLP(Attn normed ))+Attn normed , (6),

[0078] Among them, W Q ,W K ,W V is the weight parameter, Q, K, V are variables in the calculation process, Norm represents the regularization function, MLP is a multi-layer perceptron, and output is the final weighted feature representation.

[0079] Preferably, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the first aspect when executing the program.

[0080] Preferably, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.

[0081] The beneficial effects achieved by the present invention are:

[0082] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0083] First, since conventional technologies do not have effective supervision of unknown categories during training, the features of unknown categories will inevitably depend on the features of known categories, which is particularly challenging for the division of known and unknown categories. This also leads to many open-world object detection models generating false correlations when processing unknown categories. Through progressive feature fusion and attention mechanisms, the present invention enables the model to effectively fuse feature maps of different scales, enhance feature expression capabilities, and improve the detection accuracy of known categories.

[0084] Second, the present invention proposes an uncertainty estimation module, which can adjust the scores after classification to a certain extent. This method calculates the variance and mean through Monte Carlo sampling to obtain the probability that the unknown category may exist, and gives a new adjustment to the score of the unknown category in combination with the variance. This method can effectively improve the recall rate of the unknown category while maintaining the detection accuracy of the known category unchanged. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0086] Figure 1 It is a finishing flow chart of the present invention;

[0087] Figure 2 It is the overall structure diagram of the network of the present invention;

[0088] Figure 3 This is a schematic diagram of the network structure of MLLABlock attention. DETAILED DESCRIPTION

[0089] See also Figure 1 The present application discloses an open world target detection method based on progressive feature fusion and uncertainty estimation, which is characterized by comprising:

[0090] Use the trained open-world object detection model to detect natural images of natural scenes, and determine the classification results of known category features, the location information of known category features, the target scores of known category features, the classification results of unknown category features, the location information of unknown category features, and the target scores of unknown category features;

[0091] Among them, the training of the open world target detection model includes:

[0092] Step 1, obtaining an open world object detection dataset and preprocessing the data in the open world object detection dataset;

[0093] Step 2: Use the improved Fast R-CNN network to build an open-world object detection model. The open-world object detection model uses ResNet-50 as the backbone network to extract the initial multi-scale feature map; perform weighted progressive fusion on the initial multi-scale feature map to obtain a fused multi-scale feature map; perform cross-attention calculation on the fused multi-scale feature map to obtain a fixed-size feature map; perform category classification score, bounding box regression score and target score on the fixed-size feature map;

[0094] Step 3: Make an uncertainty estimate of the category classification score, use Monte Carlo sampling to calculate the variance of the category classification score as the confidence of the classification result, and quantify the uncertainty based on the confidence; judge whether the category corresponding to the category classification score belongs to a known category or an unknown category based on the uncertainty; and make a backdoor adjustment based on the uncertainty to obtain the final category classification score;

[0095] Step 4: Use the open-world object detection dataset to train the open-world object detection model.

[0096] In the embodiment of the present application, step 1, obtaining an open-world object detection dataset and preprocessing the data in the open-world object detection dataset, includes:

[0097] Step (1.1), obtaining an open-world object detection dataset, the open-world object detection dataset contains natural images with labeled object instances, and the open-world object detection dataset has multiple known categories and several unknown categories;

[0098] Step (1.2), cropping the natural image in the open-world object detection dataset to obtain an image of a preset size; modifying the instance information in the open-world object detection dataset to adapt to the image of the preset size;

[0099] Step (1.3) is to label the graph of preset size, redivide it into several tasks, and perform incremental learning on the tasks.

[0100] In the embodiment of the present application, step 2 includes:

[0101] Step (2.1), define the feature maps of different scales of the backbone network Resnet50 as {C1, C2, C3, C4, C5}, where C1 corresponds to the features of the initial convolutional layer and pooling layer, and C2, C3, C4, C5 correspond to the features of subsequent layers;

[0102] Step (2.2), substitute C1 into a convolutional layer, batch normalization layer and ReLU activation function for processing to obtain C2;

[0103] Substitute C2 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C3;

[0104] Substitute C3 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C4;

[0105] Substitute C4 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C5;

[0106] Step (2.3), use global maximum pooling for C1, reduce the width and height of C1 to 1 / 4 of the original, apply convolution with a step size of 1 and 1x1 to C2, and obtain the fused feature map P2;

[0107] Step (2.4), apply 1x1 convolution and bilinear interpolation to upsample C2 to align the spatial size of C2 with C3 to obtain C2′; apply 2x2 convolution with a step size of 2 to C3 to align the feature dimension of C3 with C2′;

[0108] Step (2.5), calculate the weight of C2′ and the weight of C3 and perform weighted fusion to obtain the fused feature map P3;

[0109] Step (2.6), repeat the above steps (2.4) and (2.5), weightedly fuse the features of C3 and C4, and the features of C4 and C5; when weighted fusing the features of C3 and C4, apply convolution with a step size of 4 and 4x4 to obtain the fused feature map P4; weighted fusing the features of C4 and C5, apply convolution with a step size of 8 and 8x8 to obtain the fused feature map P5, and obtain a set of fused multi-scale feature maps {P2, P3, P4, P5};

[0110] In step (2.7), the fused multi-scale feature map is cross-attention calculated to obtain a fixed-size feature map; the fixed-size feature map is scored for category classification, bounding box regression, and targetness.

[0111] In the embodiment of the present application, step (2.6) comprises:

[0112] For each pair of data in the feature map {(C1, C2), (C2, C3), (C3, C4), (C4, C5)}, perform the following steps:

[0113] Step (2.41), apply a convolution layer, a batch normalization layer and a ReLU activation function to a pair of feature maps respectively to obtain the first feature and the second feature;

[0114] Step (2.42), concatenate the first feature and the second feature in the channel dimension to obtain a third feature map;

[0115] Step (2.43), apply 1x1 convolution to the third feature map, and calculate the weights of feature maps at all levels;

[0116] Step (2.44), use the softmax function to normalize the weights of feature maps at all levels to obtain the fusion weight of each feature map;

[0117] Step (2.45), according to the fusion weight of each feature map, the feature maps of different scales are weightedly fused to obtain a fused feature map;

[0118] Step (2.46), apply a convolution layer, batch normalization layer and ReLU activation function to the fused feature map to obtain the final weighted fused feature map;

[0119] In step (2.47), all the final weighted fusion feature maps are combined to obtain the final weighted multi-scale fusion feature map {P2, P3, P4, P5}.

[0120] In the embodiment of the present application, in step 3, an uncertainty estimation is performed on the category classification score, and the variance of the category classification score is calculated using Monte Carlo sampling as the confidence of the classification result, and the uncertainty is quantified according to the confidence; based on the uncertainty, it is judged whether the category corresponding to the category classification score belongs to a known category or an unknown category; and a backdoor adjustment is performed based on the uncertainty to obtain the final category classification score, including:

[0121] Step (3.1), fusion feature map F for the i-th sampling i Perform N forward propagations for each pixel in, i = 1, 2, ..., N;

[0122] Step (3.2), the fusion feature map F obtained by the i-th sampling i Use classifier C to classify and get the predicted probability distribution P for each category i (c);

[0123] Step (3.3), calculate the mean μ(c) and variance σ of the predicted probability of each category c 2 (c), obtain the uncertainty; where,

[0124]

[0125] Step (3.4), based on the uncertainty, evaluate the classification score of each natural image to determine whether there is an unknown category; if the classification score of the natural image exceeds a first preset unknown classification threshold, it is determined that there is an unknown category;

[0126] In step (3.5), the calculated uncertainty is used as the uncertainty measure for backdoor adjustment.

[0127] In the embodiment of the present application, step (3) includes:

[0128] Step (3.51), perform N forward propagations on the category classification scores during the inference process of the open world object detection model;

[0129] Step (3.52), for each category c, calculate the average probability value u(c) of all N Monte Carlo samplings;

[0130] Step (3.53), for each category c, calculate the variance σ of the predicted probability of all N Monte Carlo samples 2 (c);

[0131] Step (3.54), using u(c) and σ 2 (c) evaluating the prediction uncertainty of each natural image in the open world object detection dataset for a specific category, and if the classification score of the natural image exceeds a second preset unknown classification threshold, determining that an unknown classification exists;

[0132] Step (3.55), calculate the variance σ 2 (c) As a confidence uncertainty measure for each natural image in the open-world object detection dataset to dynamically adjust the classification score and of each natural image.

[0133] In the embodiment of the present application, step (3.5) comprises:

[0134] (3.51) If the variance σ 2 (c) if it exceeds a preset threshold τ, the corresponding natural image is judged to have high uncertainty;

[0135] (3.52) The variance value σ of each high uncertainty natural image 2 (c) As an adjustment factor, it is added to the classification score of natural images;

[0136] (3.53) Calculate the classification score S′ of the natural image of each high uncertainty sample:

[0137] S′=S+β·σ 2 (c),

[0138] Among them, S is the original classification score, and β is a preset coefficient;

[0139] (3.54) If the backdoor adjusted score S′ of each category c exceeds the adjusted third classification threshold, the corresponding natural image is judged to belong to category c.

[0140] In the embodiment of the present application, step (2.7), performing cross-attention calculation on the fused multi-scale feature map to obtain a fixed-size feature map, includes:

[0141] (2.71), receiving the input feature map f, the dimension of the feature map f is (b,n,c), b is the batch size, n is the number of detection boxes, and c is the feature dimension;

[0142] (2.72), the feature dimension dim=64 and the input resolution of the feature map f is 1024;

[0143] (2.73), use the nn.Conv1d(dim, dim, 3, padding = 1, groups = dim) convolution layer to perform depth-separable convolution on the feature map f to obtain the processed feature map; add the processed feature map to the original input feature map to form a residual connection;

[0144] (2.74), normalize the input feature map, then use the linear layer to transform the feature map and apply the activation function, repeat this step twice;

[0145] (2.75), calculate the cross attention based on the input feature map input obtained in step (2.74) to obtain the weighted feature representation;

[0146] (2.75), transform the weighted feature representation through a linear layer, add the weighted feature representation and the residual connection to obtain the feature map y;

[0147] (2.76), use the nn.Conv1d(dim,dim,3,padding=1,groups=dim) convolution layer to perform depth-wise separable convolution on the feature map y;

[0148] (2.77), the input feature map output by step (2.76) is normalized and further processed through the MLP layer to generate an output feature map of fixed size.

[0149] In the embodiment of the present application, step (2.75) includes:

[0150] Expand each dimension of the obtained input feature map x into (b,n,num_heads,head_dim), where b is the batch size, n is the number of detection boxes, num_heads is the number of attention heads, and head_dim is the dimension of the attention head;

[0151] Perform self-attention calculation on x:

[0152] Q=φ(xW Q ), (1),

[0153] K=φ(xW K ), (2),

[0154] V=φ(xW V ), (3),

[0155]

[0156] Attn normed =Norm(Attn)+x, (5),

[0157] output=Norm(MLP(Attn normed ))+Attn normed , (6),

[0158] Among them, W Q ,W K ,W V is the weight parameter, Q, K, V are variables in the calculation process, Norm represents the regularization function, MLP is a multi-layer perceptron, and output is the final weighted feature representation.

[0159] In an embodiment of the present application, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the program.

[0160] In an embodiment of the present application, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0161] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0162] Those skilled in the art will readily appreciate other embodiments of the invention after considering the specification and practicing the invention invented herein. This application is intended to cover any variations, uses or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art that are not invented by the present invention. The specification and examples are intended to be exemplary only, and the true scope and spirit of the invention are indicated by the claims.

[0163] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above are only specific implementation methods of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the present application should be included in the protection scope of the present application.

Claims

1. An open-world object detection method based on progressive feature fusion and uncertainty estimation, characterized in that: include: Use the trained open-world object detection model to detect natural images of natural scenes, and determine the classification results of known category features, the location information of known category features, the target scores of known category features, the classification results of unknown category features, the location information of unknown category features, and the target scores of unknown category features; Among them, the training of the open world target detection model includes: Step 1, obtaining an open world object detection dataset and preprocessing the data in the open world object detection dataset; Step 2: Use the improved Fast R-CNN network to build an open-world object detection model. The open-world object detection model uses ResNet-50 as the backbone network to extract the initial multi-scale feature map; perform weighted progressive fusion on the initial multi-scale feature map to obtain a fused multi-scale feature map; perform cross-attention calculation on the fused multi-scale feature map to obtain a fixed-size feature map; perform category classification score, bounding box regression score and target score on the fixed-size feature map; Step 3: Make an uncertainty estimate of the category classification score, use Monte Carlo sampling to calculate the variance of the category classification score as the confidence of the classification result, and quantify the uncertainty based on the confidence; judge whether the category corresponding to the category classification score belongs to a known category or an unknown category based on the uncertainty; and make a backdoor adjustment based on the uncertainty to obtain the final category classification score; Step 4: Use the open-world object detection dataset to train the open-world object detection model.

2. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 1, characterized in that: Step 1: Obtain an open world object detection dataset and preprocess the data in the open world object detection dataset, including: Step (1.1), obtaining an open-world object detection dataset, the open-world object detection dataset contains natural images with labeled object instances, and the open-world object detection dataset has multiple known categories and several unknown categories; Step (1.2), cropping the natural image in the open-world object detection dataset to obtain an image of a preset size; modifying the instance information in the open-world object detection dataset to adapt to the image of the preset size; Step (1.3) is to label the graph of preset size, redivide it into several tasks, and perform incremental learning on the tasks.

3. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 1, characterized in that: Step 2 includes: Step (2.1), define the feature maps of different scales of the backbone network Resnet50 as {C1, C2, C3, C4, C5}, where C1 corresponds to the features of the initial convolutional layer and pooling layer, and C2, C3, C4, C5 correspond to the features of subsequent layers; Step (2.2), substitute C1 into a convolutional layer, batch normalization layer and ReLU activation function for processing to obtain C2; Substitute C2 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C3; Substitute C3 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C4; Substitute C4 into a convolutional layer, a batch normalization layer, and a ReLU activation function to obtain C5; Step (2.3), use global maximum pooling for C1, reduce the width and height of C1 to 1 / 4 of the original, apply convolution with a step size of 1 and 1x1 to C2, and obtain the fused feature map P2; Step (2.4), apply 1x1 convolution and bilinear interpolation to upsample C2 to align the spatial size of C2 with C3 to obtain C2′; apply 2x2 convolution with a step size of 2 to C3 to align the feature dimension of C3 with C2′; Step (2.5), calculate the weight of C2′ and the weight of C3 and perform weighted fusion to obtain the fused feature map P3; Step (2.6), repeat the above steps (2.4) and (2.5), weightedly fuse the features of C3 and C4, and the features of C4 and C5; when weighted fusing the features of C3 and C4, apply convolution with a step size of 4 and 4x4 to obtain the fused feature map P4; weighted fusing the features of C4 and C5, apply convolution with a step size of 8 and 8x8 to obtain the fused feature map P5, and obtain a set of fused multi-scale feature maps {P2, P3, P4, P5}; In step (2.7), the fused multi-scale feature map is cross-attention calculated to obtain a fixed-size feature map; the fixed-size feature map is scored for category classification, bounding box regression, and targetness.

4. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 3, characterized in that: Step (2.6) includes: For each pair of data in the feature map {(C1, C2), (C2, C3), (C3, C4), (C4, C5)}, perform the following steps: Step (2.41), apply a convolution layer, a batch normalization layer and a ReLU activation function to a pair of feature maps respectively to obtain the first feature and the second feature; Step (2.42), concatenate the first feature and the second feature in the channel dimension to obtain a third feature map; Step (2.43), apply 1x1 convolution to the third feature map, and calculate the weights of feature maps at all levels; Step (2.44), use the softmax function to normalize the weights of feature maps at all levels to obtain the fusion weight of each feature map; Step (2.45), according to the fusion weight of each feature map, the feature maps of different scales are weightedly fused to obtain a fused feature map; Step (2.46), apply a convolution layer, batch normalization layer and ReLU activation function to the fused feature map to obtain the final weighted fused feature map; In step (2.47), all the final weighted fusion feature maps are combined to obtain the final weighted multi-scale fusion feature map {P2, P3, P4, P5}.

5. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 1, characterized in that: In step 3, the uncertainty of the classification score is estimated, and the variance of the classification score is calculated by Monte Carlo sampling as the confidence of the classification result, and the uncertainty is quantified according to the confidence; Based on the uncertainty, it is determined whether the classification corresponding to the classification score belongs to a known category or an unknown category; The final category classification score is obtained by backdoor adjustment based on uncertainty, including: Step (3.1), fusion feature map F for the i-th sampling i Perform N forward propagations for each pixel in, i = 1, 2, ..., N; Step (3.2), the fusion feature map F obtained by the i-th sampling i Use classifier C to classify and get the predicted probability distribution P for each category i (c); Step (3.3), calculate the mean μ(c) and variance σ of the predicted probability of each category c 2 (c), obtain the uncertainty; where, Step (3.4), based on the uncertainty, evaluate the classification score of each natural image to determine whether there is an unknown category; if the classification score of the natural image exceeds a first preset unknown classification threshold, it is determined that there is an unknown category; In step (3.5), the calculated uncertainty is used as the uncertainty measure for backdoor adjustment.

6. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 5, characterized in that: Step (3) includes: Step (3.51), perform N forward propagations on the category classification scores during the inference process of the open world object detection model; Step (3.52), for each category c, calculate the average probability value u(c) of all N Monte Carlo samplings; Step (3.53), for each category c, calculate the variance σ of the predicted probability of all N Monte Carlo samples 2 (c); Step (3.54), using u(c) and σ 2 (c) evaluating the prediction uncertainty of each natural image in the open world object detection dataset for a specific category, and if the classification score of the natural image exceeds a second preset unknown classification threshold, determining that an unknown classification exists; Step (3.55), calculate the variance σ 2 (c) As a confidence uncertainty measure for each natural image in the open-world object detection dataset to dynamically adjust the classification score and of each natural image.

7. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 5, characterized in that: Step (3.5) includes: (3.51) If the variance σ 2 (c) if it exceeds a preset threshold τ, the corresponding natural image is judged to have high uncertainty; (3.52) The variance value σ of each high uncertainty natural image 2 (c) As an adjustment factor, it is added to the classification score of natural images; (3.53) Calculate the classification score S′ of the natural image of each high uncertainty sample: S′=S+β·σ 2 (c), Among them, S is the original classification score, and β is a preset coefficient; (3.54) If the backdoor adjusted score S′ of each category c exceeds the adjusted third classification threshold, the corresponding natural image is judged to belong to category c.

8. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 3, characterized in that: Step (2.7) performs cross-attention calculation on the fused multi-scale feature map to obtain a fixed-size feature map, including: (2.71), receiving the input feature map f, the dimension of the feature map f is (b,n,c), b is the batch size, n is the number of detection boxes, and c is the feature dimension; (2.72), the feature dimension dim=64 and the input resolution of the feature map f is 1024; (2.73), use the nn.Conv1d(dim, dim, 3, padding = 1, groups = dim) convolution layer to perform depth-separable convolution on the feature map f to obtain the processed feature map; add the processed feature map to the original input feature map to form a residual connection; (2.74), normalize the input feature map, then use the linear layer to transform the feature map and apply the activation function, repeat this step twice; (2.75), calculate the cross attention based on the input feature map input obtained in step (2.74) to obtain the weighted feature representation; (2.75), transform the weighted feature representation through a linear layer, add the weighted feature representation to the residual connection, and obtain the feature map y; (2.76), use the nn.Conv1d(dim,dim,3,padding=1,groups=dim) convolution layer to perform depth-wise separable convolution on the feature map y; (2.77), the input feature map output by step (2.76) is normalized and further processed through the MLP layer to generate an output feature map of fixed size.

9. The open world object detection method based on progressive feature fusion and uncertainty estimation according to claim 8, characterized in that: Step (2.75) includes: Expand each dimension of the obtained input feature map x into (b,n,num_heads,head_dim), where b is the batch size, n is the number of detection boxes, num_heads is the number of attention heads, and head_dim is the dimension of the attention head; Perform self-attention calculation on x: Q=φ(xW Q ), (1), K=φ(xW K ), (2), V=φ(xW V ), (3), Attn normed =Norm(Attn)+x, (5), output=Norm(MLP(Attn normed ))+Attn normed , (6), Among them, W Q ,W K ,W V is the weight parameter, Q, K, V are variables in the calculation process, Norm represents the regularization function, MLP is a multi-layer perceptron, and output is the final weighted feature representation.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Open world target detection method based on decoupling cascade region generation network

    CN116805389A

  • End-to-end infrared small target detection method based on Transform decoder network

    CN118505965A

  • Classification system based on multi-mode credible progressive fusion

    CN118656683A

  • Open world target detection method and system based on attention mechanism

    CN118711108A

Cited By

  • Incremental learning method for open world target detection based on large language model

    CN120747707A