Generalized small sample open set new energy power station hidden danger detection method based on HSIC
By adopting HSIC-based methods in the detection of hidden dangers of new energy power plants, a data set containing multiple categories is constructed and model weight updates are performed, and the problem of difficulty in identification in small and medium-sized samples in the existing technology is solved, and higher accuracy and generalization are achieved, especially in the recognition of unknown categories.
Patent Information
- Application Number
- CN202510307437.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
The existing methods of detecting hidden dangers in new energy power plants have shortcomings in terms of accuracy and generalization, especially in small sample scenarios, it is difficult to effectively distinguish known and unknown categories, the generalization ability is insufficient, and it is difficult to learn the decision-making boundaries of unknown categories.
A generalized small sample open-collection new energy power station potential hazard detection method is adopted based on HSIC. By constructing a data set containing basic categories, new categories and unknown categories, a backbone network, a region proposer and a region feature extractor are used to build an object detection model, and the model weight is updated and optimized through the HSIC loss function and the IoU-perceived unknown target loss function to achieve effective identification of unknown categories.
It significantly improves the rejection ability and generalization performance of the model for unknown categories, overcomes the limitations of existing methods in the detection of hidden dangers of small-sample new energy power plants, and can effectively identify unknown categories and improve the robustness and adaptability of the model.
Smart Images

Figure CN120219844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object detection, and specifically provides a generalized few-shot open-set hidden danger detection method for new energy power stations based on HSIC (Hilbert-Schmidt Independence Criterion). Background Art
[0002] Existing object detection methods usually require a large number of labeled samples for training and assume that the training set and the test set share the same categories (closed-set assumption). In real-world scenarios, the efficiency of these detectors drops rapidly when dealing with long-tailed distribution data and unknown data. Few-shot Open-set Object Detection (FOOD) is a task that combines few-shot learning with open-set recognition, mainly targeting scenarios with data imbalance and incomplete categories in the real world. For example, in the hidden danger detection of new energy power stations, the target categories are diverse and the number of trainable samples is small. It is required that the model can identify known categories trained with only a small number of samples, and at the same time, it can also identify unknown categories, which poses higher requirements for the generalization ability and the ability to handle uncertainty of the model.
[0003] However, the existing hidden danger detection methods for new energy power stations have obvious deficiencies in terms of accuracy and generalization, mainly concentrated in the following three aspects: (1) Uncertainty estimation of targets; existing methods (such as uncertainty estimation based on energy or entropy) are difficult to effectively distinguish known and unknown categories in the few-shot hidden danger detection scenario of new energy power stations and cannot construct a compact unknown decision boundary for the model. (2) Insufficient generalization ability; few-shot data easily leads to overfitting of the model to known categories, thus unable to reject unknown categories. The traditional moving weight averaging method assumes that the weight update process is linear and cannot adapt to the dynamic changes of weight updates during the training process. (3) Training data limitation; due to the absence of real unknown data, it is difficult for the model to directly learn the decision boundary of unknown categories.
[0004] Therefore, the present invention proposes a generalized few-shot open-set hidden danger detection method for new energy power stations based on HSIC, which significantly improves the generalization performance of the model while enhancing the model's ability to reject unknown categories, and overcomes the limitations of existing methods in the few-shot hidden danger detection task of new energy power stations. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a generalized few-shot open-set hidden danger detection method for new energy power stations based on HSIC.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A generalized small-sample open-set hidden danger detection method for new energy power stations based on HSIC, characterized in that the method comprises the following steps:
[0008] Step 1: Establish a data set containing different targets, and the target categories include three types: basic categories, new categories, and unknown categories;
[0009] Step 2: Construct a target detection model, including a backbone network, a region proposal network, and a region feature extractor; the image to be detected is input into the backbone network for feature extraction to obtain a feature map; the feature map is input into the region proposal network to generate candidate boxes of the target, and the feature map and the candidate boxes are input into the region feature extractor for alignment to obtain the output feature vector pair values of the candidate boxes belonging to each category;
[0010] Step 3: Calculate the evidence uncertainty of the candidate boxes belonging to each category according to Equation (1), regard the candidate boxes with high evidence uncertainty as pseudo-unknown samples, and regard the candidate boxes with low evidence uncertainty as known samples to realize the classification of the candidate boxes;
[0011]
[0012] In the formula, l k represents the output feature vector pair value of the candidate box belonging to category k, u(l k ) represents the evidence uncertainty of the candidate box belonging to category k, K represents the number of known categories, K + 2 represents the total number of categories including known categories, unknown categories, and background categories, and β k (l k ) represents the evidence intensity function of category k;
[0013] Use the classified samples to train the target detection model to obtain classification weights and regression weights; calculate the evidence uncertainty loss according to the evidence uncertainty loss function of Equation (2) during the training process;
[0014]
[0015] In the formula, λ t represents the annealing weight factor, N represents the number of samples, c i,k represents the label of sample i belonging to category k, Γ(·) represents the gamma function, represents the total evidence intensity of sample i belonging to K + 2 categories, represents the evidence intensity of sample i belonging to category k;
[0016] Calculate the unknown target loss according to the IoU-aware unknown target loss function of Equation (3);
[0017]
[0018] In the formula, b i and are the pseudo-unknown bounding box and the true bounding box respectively, represents the positioning quality parameter of the pseudo-unknown bounding box, is the probability that sample i belongs to the unknown category, is the logarithmic value of the output feature vector of sample i belonging to the unknown category, represents the logarithmic value of the output feature vector of sample i belonging to category k, is the logarithmic value of the output feature vector of sample i belonging to the true category, and λ is a constant;
[0019] Store the classification weights and regression weights in the classification weight memory bank and the regression weight memory bank respectively. During the training process, update the classification weight memory bank and the regression weight memory bank every S iterations. The update process is as follows:
[0020] Calculate the average classification weight and the average regression weight, calculate the HSIC values of the classification weights and the regression weights using the HSIC loss, and update the classification weights and the regression weights according to equations (5) and (6);
[0021]
[0022] In the formula, are the updated classification weights and regression weights respectively, h cls and h reg are the HSIC values of the classification weights and the regression weights respectively, θ cls and θ reg are the current classification weights and the current regression weights respectively, are the average classification weight and the average regression weight respectively;
[0023] Add the updated classification weights and regression weights to the classification weight memory bank and the regression weight memory bank respectively, and remove the oldest classification weights and regression weights until the training is completed to obtain the trained object detection model;
[0024] Step 4: Use the trained object detection model for hidden danger detection in new energy power stations.
[0025] Furthermore, the weight update is constrained according to the following HSIC loss function:
[0026]
[0027] In the formula, and represent the expected values of the classification weights and the regression weights, represents the HSIC loss between the current classification weight and the average classification weight, Represents the HSIC loss between the current regression weight and the average regression weight.
[0028] Furthermore, the training loss is calculated according to the following loss function:
[0029] L = L rpn + L ce + L reg + λ1L EDL + λ2L U + λ3L HSIC (8)
[0030] In the formula, L ce is the cross-entropy loss function, L reg is the regression loss function, is the loss function of the region proposal network, M is the number of candidate boxes generated by the region proposal network, p j is the true label of the j-th candidate box, is the probability that the model predicts the j-th candidate box as the foreground, t j 、 are the true bounding box coordinates and predicted bounding box coordinates of the j-th candidate box respectively, smooth(·) is the smooth loss function, and λ1, λ2 and λ3 are weight coefficients.
[0031] Furthermore, the backbone network and the region feature extractor adopt the ResNet series network, and the region proposal network adopts the Faster R-CNN network.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] 1. The present invention integrates weight update based on HSIC, unknown target optimization based on IoU awareness, and pseudo-unknown sample mining of evidence deep learning, enabling the model to not only learn well on limited known-class data, significantly improving the generalization performance of the model, overcoming the limitations of existing methods in the few-shot hidden danger detection task of new energy power stations, but also effectively coping with the challenges of unknown-class data and identifying unknown classes.
[0034] 2. In the face of few-shot training, a moving weight averaging strategy based on the Hilbert-Schmidt independence criterion (HMWA) is proposed. Different from the traditional moving weight averaging method (MWA), this method uses the Hilbert-Schmidt independence criterion to measure the independence between the current weight and the average weight, is more adaptable to non-stationary data, flexibly adjusts the weight update method according to the weight update strategy during the training process, thereby improving the generalization ability of the model on unknown-class data, avoiding the linear assumption of the weight update process of the traditional MWA method, and enhancing the robustness and adaptability of the model.
[0035] 3. For the detection of hidden dangers of unknown categories, an IoU-aware unknown target optimization method is proposed. Existing methods usually ignore the localization differences between known and unknown category targets, which may lead to misclassifying known category targets as unknown category targets. Therefore, the method of the present invention introduces IoU-aware unknown target optimization, calculates the probability that a target belongs to an unknown category according to the overlap degree (IoU) between the predicted box and the ground truth box, and avoids misjudgment due to close positions, so as to better separate known and unknown category targets.
[0036] 4. For the mining of pseudo-unknown samples, an improved unknown sample mining method is proposed. Traditional methods use metrics such as energy scores to select pseudo-unknown samples, but may not be able to accurately capture the true situation of uncertainty. Therefore, this method introduces uncertainty estimation based on evidence deep learning (EDL), models the probability using the Dirichlet distribution, and thus more accurately mines pseudo-unknown samples and improves the model's processing ability for unknown categories. Description of the Drawings
[0037] Figure 1 is the overall flowchart of the present invention;
[0038] Figure 2 is the detection result diagram of the present invention method and the OpenDet method for photovoltaic modules;
[0039] Figure 3 is the detection result diagram of the present invention method and the OpenDet method for wind turbine blades. Detailed Embodiment
[0040] The technical solutions of the present invention will be introduced in detail below in conjunction with the drawings and specific embodiments, but the protection scope of this application is not limited thereto.
[0041] The present invention provides a generalized small-sample open-set hidden danger detection method for new energy power stations based on HSIC (hereinafter referred to as the method, see Figures 1 to 3 ), including the following steps:
[0042] The first step: Establish a data set containing different targets, and the target categories include three types: basic categories, new categories, and unknown categories;
[0043] The goal of small-sample open-set object detection is to train a detection model under the open-set assumption so that it can correctly classify objects of K + 2 categories during the test phase, including K known categories (including basic categories and new categories), an unknown category, and a background category.
[0044] Given a data set D = {(x, y)}, x = {x1, x2,..., x i ,...} represents samples, Denote the label as c i 、 denote the true class and true bounding box of sample i; the dataset D contains the training set D tr and the test set D te ; the training set D tr contains K known classes, each with M-shot support samples, and the set of known classes C K = C B ∪C N = {1, 2,..., K}, C B = {1, 2,..., B} represents the set of base classes, B represents the number of base classes, and C N = {B + 1, B + 2,..., B + P} represents the set of new classes, P represents the number of new classes, and K = B + P. The test set D te is used to evaluate the detector and contains known and unknown classes, with no overlap between the known and unknown class labels.
[0045] Step 2: Construct an object detection model, use the object detection model to generate candidate boxes, and obtain the output feature vector pair values of the candidate boxes belonging to each class;
[0046] As Figure 1 shown, the object detection model includes a backbone network, a region proposal network, and a region feature extractor; the backbone network is used to extract the feature map of the input image, input the feature map into the region proposal network composed of multiple convolutional layers to generate candidate boxes, and multiple candidate boxes are generated for each object; input the feature map obtained by the backbone network and the candidate boxes generated by the region proposal network into the region feature extractor for alignment to obtain the output feature vector pair values of the candidate boxes belonging to each class. The backbone network and the region feature extractor can adopt the ResNet series network, and the region proposal network can adopt the Faster R-CNN network.
[0047] Step 3: Process the output feature vector pair values of the candidate boxes using the evidence uncertainty estimated by the Dirichlet probability distribution, calculate the evidence uncertainty of the candidate box belonging to class k according to Equation (1), regard the candidate boxes with high evidence uncertainty as pseudo-unknown samples, and the candidate boxes with low evidence uncertainty as known samples to achieve candidate box classification;
[0048]
[0049] In the formula, l k represents the output feature vector pair value of the candidate box belonging to class k, and u(l k ) represents the evidence uncertainty of the candidate box belonging to class k; β k (l k ) represents the evidence strength function of class k, βk (l k ) = e k +1, e k represents the evidence of class k, e k = exp(l k ), exp(·) represents the exponential activation function;
[0050] The target detection model is trained using the training set, and the classification weights and regression weights of the model are updated based on the moving weight averaging strategy of HSIC; during the training process, the evidence uncertainty loss is calculated according to the evidence uncertainty loss function in Equation (2) for optimizing the uncertainty estimation of the target;
[0051]
[0052] In the formula, N represents the number of samples, c i,k represents the label that sample i belongs to class k, Γ(·) represents the gamma function, represents the total evidence strength that sample i belongs to K + 2 classes, represents the evidence strength that sample i belongs to class k; λ t = n c exp{-(lnn c / T)t} ∈ [n c , 1] represents the annealing weight factor, n c is a positive constant much less than 1, t is the current training iteration number, and T is the total training iteration number.
[0053] Considering the overlap between the pseudo-unknown bounding box b i and the ground truth bounding box as well as the probability that the sample belongs to the unknown class, in order to accurately distinguish unknown targets, an optimization method for unknown targets based on IoU (Intersection over Union) perception is proposed, so as to use the intersection over union to dynamically adjust the perception ability of the optimization model for unknown targets. Then, the loss function for unknown targets based on IoU perception is:
[0054]
[0055]
[0056] In the formula, represents the localization quality parameter of the pseudo-unknown bounding box, is the probability that sample i belongs to the unknown class, is the logarithm value of the output feature vector of sample i belonging to the unknown class, represents the logarithm value of the output feature vector of sample i belonging to class k, is the logarithmic value of the output feature vector of sample i belonging to the true category, and λ is a non - negative constant with a small value;
[0057] In the IoU - aware unknown object loss L U During the optimization process, when selecting pseudo - unknown bounding boxes from the foreground or background, a high unknown probability should correspond to a lower IoU score; when the IoU of the pseudo - unknown bounding box ≥ 0.5, the pseudo - unknown bounding box is classified as the foreground; when 0.5 > IoU ≥ 0.05, it is classified as the background. In fact, when the IoU of the foreground and background is close to or equal to 0.5 and 0 respectively, these pseudo - unknown bounding boxes are considered to be more uncertain. The IoU of pseudo - unknown bounding boxes from the background and foreground is mainly distributed in 0 - 0.1 and 0.5 - 0.6, which is consistent with the optimization purpose of the model, that is, to find high - uncertainty samples with low IoU scores. For pseudo - unknown bounding boxes from the foreground, is close to 0.5, while for pseudo - unknown bounding boxes from the background, is close to λ. Therefore, in the small - sample open - set object detection task, the impact of IoU on the uncertainty of candidate boxes is crucial. By considering the IoU between pseudo - unknown bounding boxes and the true bounding boxes of known classes, a better object function can be designed. The constants 1 and 0.5 in Equation (4) are used to balance and positive - value the IoU - aware weights of the foreground and background.
[0058] After training, the classification weights and regression weights of the model are obtained; the classification weights and regression weights are stored in the classification weight memory bank Ω cls and the regression weight memory bank Ω reg respectively. These weights need to be updated as the number of model training iterations increases; during the training process, the classification weight memory bank and the regression weight memory bank are updated every iteration; the average classification weight and the average regression weight are calculated. The HSIC loss is used to measure the independence between the current classification weight θ cls and the average classification weight and between the current regression weight θ reg and the average regression weight Then, the HSIC value of the classification weight and the HSIC value of the regression weight The classification weights and regression weights are updated according to the HSIC loss, the current weights, and the average weights. The update formulas are:
[0059]
[0060] In the formula, are the updated classification weights and regression weights respectively;
[0061] Add the updated classification weights to the classification weight memory bank and remove the oldest classification weights, add the updated regression weights to the regression weight memory bank and remove the oldest regression weights; the weight update is constrained by the following HSIC loss function:
[0062]
[0063] In the formula, and represent the expected values of the classification weights and regression weights; if the HSIC value is low, it indicates that the independence of the current weights from the average weights is strong (the correlation is weak), and the weight learning still needs further optimization, and the HSIC loss is large at this time; if the HSIC value is high, it indicates that the independence of the current weights from the average weights is weak (the correlation is strong), and the current weights have fully learned the important samples, and the HSIC loss is small.
[0064] Update the weights of the model using HSIC. Specifically, by calculating the independence between the current weights and the average weights, this metric is used to update the weights of the model. HSIC is a statistical measure of the independence between two random variables. By using HSIC to measure the independence between the current weights and the average weights, it is possible to identify when the current weights are significantly different from the previous weights, indicating that the model has generated new important data and weight update is required. Compared with the traditional weight moving average method that uses the weighted average of the current weights and the previous weights to update the model weights, the HSIC-based moving weight average strategy is more suitable for non-stationary data and does not rely on data that changes linearly over time, thus providing a more flexible and adaptive method to update the model weights and enabling it to better adapt to the distribution of weight data over time.
[0065] Furthermore, train the model in an end-to-end manner by minimizing the following loss function:
[0066] L = L rpn + L ce + L reg + λ1L EDL + λ2L U + λ3L HSIC (8)
[0067] In the formula, L ce is the cross-entropy loss function, L reg is the regression loss function; is the loss function of the region proposal generator, M is the number of candidate boxes generated by the region proposal generator, p j is the true label of the j-th candidate box, is the probability that the model predicts the j-th candidate box as the foreground, t j 、 They are the true bounding box coordinates and predicted bounding box coordinates of the j-th candidate box respectively, smooth(·) is the smooth loss function, and λ1, λ2, and λ3 are weight coefficients;
[0068] The definition of the smooth loss function is as follows:
[0069]
[0070] Step 4: Use the trained object detection model for hidden danger detection in new energy power stations.
[0071] Embodiment
[0072] To verify the effectiveness of the method of the present invention, different experimental data sets are designed, including VOC10-5-5 data set, VOC-COCO data set, and LVIS315-454-461 data set.
[0073] VOC10-5-5 data set: Divide the 20 categories of the publicly available PASCAL VOC data set into 10 basic categories, 5 new categories, and 5 unknown categories. For each new category, 1, 3, 5, and 10 objects are sampled from the training sets and validation sets of VOC07 and VOC12 respectively for few-shot training, and the test set of VOC07 is used as the test data.
[0074] VOC-COCO data set: Use the training sets and validation sets of PASCAL VOC as the training data for the basic categories. Select 20 categories that do not overlap with the 20 categories of VOC from the training set of MSCOCO2017 as new categories and sample 1, 5, 10, and 30 objects from them respectively, and the remaining 40 categories are used as unknown categories; use the val2017 set of COCO as the test data.
[0075] LVIS315-454-461 data set: This data set has a natural long-tail distribution. The categories are divided into 315 high-frequency categories (appearing in more than 100 images), 461 common categories (appearing in 10 to 100 images), and 454 rare categories (appearing in less than 10 images). The high-frequency categories, rare categories, and common categories are used as the basic categories, new categories, and unknown categories respectively. The rare categories can be directly used as few-shot training data (without manual division of M-shot), and the validation set of LVIS is used as the test data.
[0076] Compare the method of the present invention with common few-shot open-set object detection methods. The comparison methods include DS [1] , PROSER [2] , OPENDET [3]。The method of the present invention and all comparative methods use the same backbone network (Resnet50), region proposal generator, and region feature extractor. One, three, five, and ten samples of each known class are selected for training, and mAP is used K 、mAP N to measure the detection performance of the model for known classes and new classes in known classes. WI is used to measure the ability of the model to reject unknown class samples in the open-set scenario. The smaller the value, the stronger the ability of the model to reject unknown class samples. AOSE is used to measure the uncertainty information output by the model. By analyzing the entropy value of unknown class samples on the Softmax distribution, the discrimination ability of the model for unknown classes can be quantified. The smaller the value, the higher the entropy value of the model for unknown class samples, the stronger the prediction uncertainty of the model for unknown classes, and the better the open-set recognition ability. Recall rate AR U is used to measure the recall performance of the model for unknown class samples. The larger the value, the better the performance of identifying unknown class samples. The test results of different methods on different datasets are shown in Tables 1-3.
[0077] Table 1 Statistical results of test on VOC10-5-5 dataset for different methods
[0078]
[0079] Table 2 Statistical results of test on VOC-COCO dataset for different methods
[0080]
[0081] Table 3 Statistical results of test on LVIS315-454-461 dataset for different methods
[0082] LVIS315-454-461 <![CDATA[mAP K ↑(%)]]> <![CDATA[mAP N ↑(%)]]> WI↓(%) AOSE↓ <![CDATA[AR U ↑(%)]]> DS 17.3 12.8 15.5 2107.5 1.4 PROSER 17.1 12.5 17.8 2158.8 2.7 OPENDET 17.2 12.4 14.8 1923.5 3.1 the present invention 20.9 16.8 7.9 1200.2 9.1
[0083] As can be seen from Table 1, the method of the present invention is superior to the existing methods in rejecting unknown classes and recognizing known classes on the same type of dataset. Specifically, when the method of the present invention is trained with one, three, five, and ten samples, mAP K / mAP N reaches 45.8% / 18.3%, 49.3% / 30.3%, 53.8% / 34.7%, 57.2% / 41.3%, and the average value is 51.5% / 31.5%, exceeding the existing methods by 4.4% / 14.3%. WI / AOSE are 4.3% / 367.2, 4.3% / 600.6, 4.2% / 719.6, 4.1% / 818.4, and the average value is 4.2% / 626.4, with an optimization of 4.8% / 313.3. Recall rate AR UThey are 31.4%, 32.5%, 33.0%, and 33.2% respectively, with an average value of 32.5%, representing a 15.3% increase.
[0084] As can be seen from Table 2, the method of the present invention is superior to the existing methods in rejecting unknown classes and recognizing known classes on different types of datasets. Specifically, when the method of the present invention is trained with 1, 5, 10, and 30 samples, mAPK / mAPN reaches 18.8% / 4.5%, 20.4% / 12.4%, 22.8% / 14.7%, and 24.3% / 18.6% respectively, with an average value of 21.6% / 12.6%, exceeding the existing methods by 2.8% / 4.3%; WI / AOSE are 4.6% / 319.2, 5.4% / 309.9, 6.1% / 385.2, and 5.0% / 317.2 respectively, with an average value of 5.3% / 332.9, optimizing by 4.3% / 870.3; the recall rate AR U They are 14.2%, 15.7%, 16.7%, and 17.6% respectively, with an average value of 16.1%, representing a 10.2% increase.
[0085] As can be seen from Table 3, the method of the present invention is superior to the existing methods in rejecting unknown classes and recognizing known classes on datasets with the characteristics of a wide range of categories and long-tail distribution. Specifically, the mAP K of the method of the present invention reaches 20.9%, representing a 3.6% increase; the mAP N reaches 16.8%, representing a 4% increase; WI reaches 7.9%, representing a 6.9% decrease; AOSE reaches 1200.2, optimizing by 723.3; the recall rate ARU is 9.1%, representing a 6% increase.
[0086] Experimental data for hidden danger detection in new energy power stations: There are 3520 photovoltaic infrared detection images in the new energy power station. Regarding hot spots as known classes and abnormally shaped spots as unknown classes; there are 1198 visible light and infrared light images of the fan blades in the new energy power station. Regarding six common types of defects as known classes, including oil leakage, dirt adhesion, paint peeling, rusting, fading, and scratches, and setting dirt adhesion as the unknown class during the test process. Figure 2 、 3 They are the detection result diagrams of photovoltaic modules and fan blades by OpenDet and the method of the present invention respectively. As can be seen from the diagrams, the method of the present invention has good unknown sample recognition ability.
[0087] In summary, the method of the present invention can perform good class recognition and unknown class prediction on datasets with various characteristics. This is because the method of the present invention mines pseudo-unknown samples and updates the model weights using the HSIC-based moving weight averaging strategy, improving the stability and generalization of small-sample open-set object detection.
[0088] Source of the above existing method:
[0089] [1]Dimity Miller,Lachlan Nicholson,Feras Dayoub,and Niko Sunderhauf.2018.Dropout sampling for robust object detection in open-set conditions.In Proceedings of the IEEE Int.Conf.Robot.Autom.(ICRA).3243–3249.
[0090] [2]Dawei Zhou,Hanjia Ye,and Dechuan Zhan.2021.Learning placeholders for open-set recognition.In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR).4401–4410.
[0091] [3]Jiaming Han,Yuqiang Ren,Jian Ding,Xingjia Pan,Ke Yan,and Guisong Xia.2022.Expanding Low-Density Latent Regions for Open-Set Object Detection.In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR).
[0092] What is not described in the present invention applies to the prior art.
Claims
1. A generalized small sample open set new energy power station hidden danger detection method based on HSIC, characterized in that: The method comprises the following steps: Step 1: Create a data set containing different targets. The target categories include basic categories, new categories, and unknown categories. Step 2: Build a target detection model, including a backbone network, a region proposer, and a region feature extractor; input the image to be detected into the backbone network for feature extraction to obtain a feature map; input the feature map into the region proposer to generate a candidate box for the target, input the feature map and the candidate box into the region feature extractor for alignment, and obtain the output feature vector logarithm of the candidate box belonging to each category; Step 3: Calculate the uncertainty of the evidence that the candidate box belongs to each category according to formula (1), regard the candidate box with high uncertainty as a pseudo unknown sample, and regard the candidate box with low uncertainty as a known sample, so as to realize the classification of the candidate box; In the formula, l k Indicates that the candidate box belongs to the output feature vector logarithm of category k, u(l k ) represents the uncertainty of the evidence that the candidate box belongs to category k, K represents the number of known categories, K+2 represents the total number of categories including known categories, unknown categories and background analogs, β k (l k ) represents the evidence strength function of category k; Use the data set to train the target detection model to obtain classification weights and regression weights; During the training process, the evidence uncertainty loss is calculated according to the evidence uncertainty loss function of formula (2); In the formula, λ t represents the annealing weight factor, N represents the number of samples, c i,k represents the label of sample i belonging to category k, Γ(·) represents the gamma function, Indicates the total evidence strength that sample i belongs to K+2 categories, Indicates the strength of evidence that sample i belongs to category k; The unknown target loss is calculated according to the IoU-aware unknown target loss function of formula (3); Where b i , They are the pseudo unknown bounding box and the real bounding box, represents the localization quality parameter of the pseudo unknown bounding box, is the probability that sample i belongs to the unknown category, is the logarithm of the output feature vector of sample i belonging to the unknown category, The logarithm of the output feature vector indicating that sample i belongs to category k, is the logarithm of the output feature vector of the true category of sample i, and λ is a constant; The classification weight and regression weight are stored in the classification weight memory and regression weight memory respectively. During the training process, the classification weight memory and regression weight memory are updated every S iterations. The update process is: Calculate the average classification weight and average regression weight, use HSIC loss to calculate the HSIC value of the classification weight and regression weight, and update the classification weight and regression weight according to equations (5) and (6); In the formula, are the updated classification weight and regression weight, respectively, cls 、h reg are the HSIC values of classification weight and regression weight, θ cls ,θ reg are the current classification weight and the current regression weight, respectively. are the average classification weight and the average regression weight respectively; The updated classification weights and regression weights are added to the classification weight memory bank and the regression weight memory bank respectively, and the oldest classification weights and regression weights are removed until the training is completed to obtain the trained object detection model; Step 4: Use the trained target detection model for hidden danger detection in new energy power stations.
2. The generalized small sample open set new energy power station hidden danger detection method based on HSIC according to claim 1 is characterized in that: The weight update is constrained according to the following HSIC loss function: In the formula, represents the expected value of classification weight and regression weight, Represents the HSIC loss of the current classification weight and the average classification weight, Represents the HSIC loss of the current regression weight and the average regression weight.
3. The generalized small sample open set new energy power station hidden danger detection method based on HSIC according to claim 1 is characterized in that: The training loss is calculated according to the following loss function: L=L rpn +L ce +L reg +λ1L EDL +λ2L U +λ3L HSIC (8) Where, L ce is the cross entropy loss function, L reg is the regression loss function, is the loss function of the region proposer, M is the number of candidate boxes generated by the region proposer, p j is the true label of the jth candidate box, is the probability that the model predicts the jth candidate box as the foreground, t j , are the true bounding box coordinates and predicted bounding box coordinates of the jth candidate box, respectively, smooth(·) is the smoothing loss function, and λ1, λ2, and λ3 are weight coefficients.
4. The generalized small sample open set new energy power station hidden danger detection method based on HSIC according to any one of claims 1 to 3, characterized in that: The backbone network and the regional feature extractor adopt the ResNet series network, and the regional proposer adopts the Faster R-CNN network.
Citation Information
Patent Citations
Small sample target detection method based on self-supervised contrast constraint
CN114841257A