New energy power station inspection method based on multi-mode large model small sample open set
By combining a large multimodal model with image and text information, a generalized small-sample open-set target detection model is constructed, which solves the problems of incomplete information and category confusion in the inspection of new energy power stations, and achieves efficient unknown category recognition and fine-grained decision-making.
Patent Information
- Application Number
- CN202510829255.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing inspection methods for new energy power plants are inefficient when faced with rare defects. Single-modal image modeling information is incomplete, and the model easily misclassifies unknown categories as known categories, leading to category confusion and making it difficult to form effective zero-shot class discrimination.
A multimodal large model is used to combine image and text information to build a generalized small-sample open-set target detection model. The model decision boundary is optimized by mining pseudo samples of unknown classes and a cost-aware loss function to form fine-grained category decisions.
It achieves the goal of improving the detection accuracy of new energy power station inspections under small sample conditions, alleviates the problems of insufficient information and category confusion in single-modal models, and improves the ability to recognize unknown categories.
Smart Images

Figure CN120747665A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal large-model new energy power station inspection, and specifically is a new energy power station inspection method based on a multimodal large-model small sample open set. Background Art
[0002] As renewable energy power plants continue to expand in scale, traditional manual inspection methods are becoming inefficient and unable to meet the inspection needs of large-scale plants. Existing inspection methods for renewable energy power plants primarily rely on drones equipped with infrared thermal imaging cameras to capture infrared and visible light images. These images are then used to identify defects using deep learning-based object detection. Intelligent inspection technologies like drones can quickly cover large areas and improve inspection efficiency.
[0003] Methods based on deep learning rely on a large amount of manually annotated data, which requires a lot of data preparation and has low algorithm execution efficiency. However, in the inspection of new energy power plants, some defects or anomalies may be very rare, resulting in a limited number of samples. Multimodal large models can effectively learn in small sample situations and enhance the model's understanding of new concepts through cross-modal information, which is of great significance for handling rare events in power plants. Lv et al. (Lv Tiangen, Hong Richang, He Jun. Multimodal guided local feature selection small sample learning method [J]. Journal of Software, 2023, 34(05): 2068-2082.) proposed a method based on multimodal text feature measurement. By locally aligning text features with image features, the prototype measurement method is used to detect small sample targets in the image. However, the above method is not effective when facing open set tasks. Although the target detection method based on large models has made great progress in the field of zero-shot learning, it still faces a serious open set problem. The model is prone to detect unknown classes as known classes with text hints with high confidence. Generalized small-sample open set learning refers to using a small number of training samples to enable the model to detect known classes in small samples and known classes in zero samples, while rejecting interference from unknown classes. Currently, large models still face the following two challenges in small-sample learning:
[0004] (1) Unimodal image modeling suffers from incomplete information: Unimodal image modeling relies solely on visual information and lacks support from other modal data (such as text, sound, etc.). This limits the model’s comprehensive understanding of the scene. When the model has never seen an instance of a specific category, it is difficult to form a discriminative decision boundary for the zero-shot class, resulting in difficulty in zero-shot recognition.
[0005] (2) Open-set category confusion: The model overfits the text prompts and image training data of known classes, and loses the ability to generalize the representation of unknown classes. When the model has never seen the category prompt corresponding to the target, the model will make an incorrect prediction and detect the unknown class object as a known class with similar appearance, causing category confusion. Summary of the Invention
[0006] To address the shortcomings of existing technologies, the present invention proposes a method for inspecting new energy power plants using a multimodal large model and a small open-set sample size. This method, based on multimodal modeling, jointly models text and images to construct a generalized small open-set target detection model. Furthermore, it incorporates a bipartite graph matching-based method for mining pseudo-samples of unknown classes, multivariate text learnable hints, and a cost-aware unknown class loss function. This allows the detection model to form a more fine-grained category decision boundary, resulting in higher detection accuracy.
[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0008] A new energy power station inspection method based on a multimodal large model and a small sample open set, characterized in that the method comprises the following steps:
[0009] Step 1: Obtain photovoltaic inspection image sets, wind turbine inspection image sets, and tower inspection image sets from the inspection of new energy power stations. Each image set contains corresponding defect images and non-defect images. The number of images in each image set is no less than 100. Each image data in each image set includes an image with a corresponding defect category label and a corresponding text description. The defect images in each image set cover all defect categories currently known to the corresponding subject.
[0010] Step 2: Build a small sample open set detection model based on a multimodal large model
[0011] The detection model includes an end-to-end ViT model, a text processing module, and an unknown class pseudo sample mining module. The end-to-end ViT model consists of three parts: an image feature embedding module, a Transformer encoder, and an MLP classification head. The end-to-end ViT model is the basic detector.
[0012] The image feature embedding module first divides the input image P0 into multiple small blocks of equal size, then adds a position code corresponding to its position in the image P0 to each small block, and then flattens each small block into a one-dimensional vector. These flattened small blocks are then embedded into a high-dimensional feature vector through a convolutional layer or linear transformation. Finally, a learnable category token is added to the front of the high-dimensional feature vector as a representative of global information, which is finally used for classification.
[0013] The Transformer encoder contains multiple encoder layers. The output of the previous encoder layer serves as the input of the next encoder layer. Each encoder layer includes a first normalization layer, a multi-head attention layer, a second normalization layer, and a feedforward neural network. The feature T0 input to an encoder layer is first processed by the first normalization layer, and the result is input to the multi-head attention layer. The output of the multi-head attention layer is residually connected with the feature T0. The obtained feature T1 is processed by the second normalization layer and then input into the feedforward neural network. The output of the feedforward neural network is residually connected with the feature T1, and the obtained feature T2 is used as the output of the encoder layer.
[0014] The output category token of the last encoder layer of the Transformer encoder is used as the input of the MLP classification head. The MLP classification head classifies the input and obtains the final prediction result.
[0015] The text processing module includes a text encoder in the pre-trained CLIP model and a multi-text science department embedding query library. First, the single text prompts of all possible results of the new energy power station inspection are modified into multi-text learnable prompts. Then, the text encoder is used to extract the feature vector of each learnable prompt. Finally, each learnable prompt and the corresponding feature vector are embedded into a piece of data for storage, thereby obtaining the multi-text science department embedding query library.
[0016] Image P0 is processed by the image feature embedding module and the Transformer encoder in turn to obtain the feature T p , feature T p The image feature f is obtained through linear projection. The category similarity between the image feature f and the multi-text science embedding query library is calculated, and then a maximum pooling operation is performed to obtain the category prediction similarity score S. The dimension of S is B×K, where B represents the number of prediction boxes and K is the total number of prediction categories set, which is the sum of the number of all known categories and the number of set unknown categories.
[0017]
[0018] sin(,) represents cosine similarity, τ represents temperature coefficient, which is a set non-zero constant, q i represents the feature vector of the i-th learnable hint in the multi-text science department embedding query base, and Z is the number of learnable hints in the multi-text science department embedding query base;
[0019] The larger the element value in S is, the more similar the corresponding prediction box is to the corresponding learnable hint, that is, the learnable hint is the predicted label of the prediction box. According to S, the predicted label of each prediction box of the image feature f is obtained;
[0020] The feature T output by the Transformer encoder p , is input to the MLP classification head, which outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each class label; the N prediction boxes are input to the unknown class pseudo sample mining module; the unknown class pseudo sample mining module first constructs a cost matrix containing virtual unknown classes by adding unknown class virtual nodes to the bipartite graph, and then uses the Kuhn-Munkres algorithm to search for the best match of N elements from the cost matrix with the lowest cost Then calculate the cost-aware unknown class loss, the cost-aware unknown class loss function L ux It is defined as formula (6):
[0021]
[0022] Indicates the best match The intersection of the unknown class prediction box and the known class real box in is: Indicates the best match The predicted probability of the unknown class in c i is the target class label, are the coordinates of the center of gravity of the ground truth box and its height and width relative to the image size, For the best match The coordinates of the center of gravity of the predicted box in and its height and width relative to the image size; For the best match The index labels in The predicted label is category c i The predicted probability of the corresponding category label of each prediction box is output by the MLP classification head;
[0023] Step 3: Train a small sample open set detection model based on a multimodal large model
[0024] Step 3.1 First, adjust the format of each image in the photovoltaic inspection image set, wind turbine inspection image set, and tower inspection image set in step 1 to meet the input format requirements of the basic detector, and then divide the defect types in each image set into three parts, the first part of which is small sample known class data, the second part is zero sample known class data, and the third part is unknown class data; small sample known class data means that the corresponding defect category on the image has been labeled, and the text description of the image covers the corresponding defect category; zero sample known class data means that the corresponding defect category on the image is not labeled, but the text description of the image covers the corresponding defect category; unknown class data means that the corresponding defect category on the image is labeled as "unknown", and the text description of the image covers the "unknown" defect category;
[0025] In step 3.2, the base detector is the pre-trained ViT-B32 or ViT-L14 released by Google DeepMind; set the training batch size to 1, set the learning rate, hyperparameters, and maximum iteration rounds;
[0026] Step 3.3: The image data of the photovoltaic inspection image set, the wind turbine inspection image set and the tower inspection image set are input into the pre-trained basic detector in step 3.2 respectively. An image data includes image P x and the corresponding text description W x , image P x After being processed by the image feature embedding module and the Transformer encoder, the output feature T x , feature T x After linear projection, the image feature f is obtained x ;Describe the text W x Input into the text processing module, first through the text encoder extraction, get the feature vector of the text description, then the text description W x And the corresponding feature vector is stored in the multi-text science embedded query library; then the image feature f is calculated x The category similarity with the multi-text science department embedding query library is subjected to the maximum pooling operation to obtain the category prediction similarity score S x , get the image feature f x The predicted label for each predicted box;
[0027] At the same time, the feature T x The input is to the MLP classification head, which outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each prediction label. The prediction label corresponding to the maximum probability is the category label of the prediction box; the N prediction boxes are input to the unknown class pseudo sample mining module; the unknown class pseudo sample mining module first constructs a cost matrix containing virtual unknown classes by adding virtual nodes of unknown classes in the bipartite graph, and then uses the Kuhn-Munkres algorithm to search for the best match of N elements from the cost matrix at the lowest cost, and calculates the cost-aware unknown class loss function L ux The value of the unknown class loss function L ux The value of the gradient backpropagation method is used to embed the text description W in the query library of the multi-text science department x The value of the corresponding eigenvector is corrected;
[0028] According to the image P x The marked target box and the marked defect category label, category prediction similarity score S x, the MLP classification head outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each prediction label and the unknown class loss function L ux , calculate the image P x The training loss L is:
[0029]
[0030] Among them, L ce (c i ) represents the best match in the unknown class pseudo sample mining module The cross entropy classification loss is , where the best match The true category label is the label of the input image itself, and its predicted label is the similarity score S predicted according to the category x income; represents the sum of the IoU loss and L1 regression loss between the defect location prediction box output by the MLP classification head and the true box, λ iou and λ L1 is the weighting coefficient; ω is the unknown class loss function L ux The weighting coefficient of
[0031] According to the image P x The training loss value is used to minimize the training loss. The AdamW optimizer is selected to update the network parameters in the basic detector once to complete the training of one image data. Then the next image data is input into the basic detector. The network parameters of the basic detector when the previous image data is completed are used as the initial parameters for the training of the next image data. The training process of one image data is repeated continuously until the last image data is trained and a round of training is completed. The network parameters of the basic detector that completed the previous round of training are used as the initial parameters for the next round of training. The training process is repeated continuously until the training round reaches the preset value and the basic detector with the optimal parameters is obtained.
[0032] Step 4: New Energy Power Station Inspection
[0033] Obtain inspection photos of photovoltaic, wind turbines, or towers taken at new energy power plants. First, adjust the format of the photos to meet the input format requirements of the basic detector. Then, input them into the basic detector with the optimal parameters in step 3. The MLP classification head outputs the detection results, determining whether the corresponding inspected components contain defects and the defect category.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention provides a new energy power station inspection method based on a multimodal large model and a small sample open set, constructs a multi-text learnable prompt, and guides the detection model to form a more fine-grained category decision boundary by aggregating task-related text information, thereby alleviating the problem that a single text prompt causes the model to produce an overly broad decision boundary. Due to the lack of real unknown class samples during the training process, the method of the present invention regards the mining of unknown class pseudo samples as a bipartite graph matching task for the first time, and constructs a cost matrix by adding unknown class virtual nodes, mining unknown class pseudo samples, and then optimizing the model to form a compact unknown class decision boundary. In response to the problem that known classes and unknown classes in small samples are easily confused, the method of the present invention proposes a cost-aware unknown class optimization loss, and improves the open set detection performance of the model by considering the classification and positioning quality of unknown classes. The method of the present invention realizes the transition from a small unimodal visual model to a large multimodal visual model, and only requires a small amount of training data to obtain better generalized small sample open set target detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of the principle and structure of a detection model of a new energy power station inspection method based on a multimodal large model and small sample open set in the present invention.
[0036] Figure 2 This is a schematic diagram of the principle of a bipartite graph matching algorithm for the unknown class pseudo sample mining module of a detection model based on a multimodal large model and small sample open set new energy power station inspection method, where (a) is a schematic diagram of the bipartite graph matching process of the real box and the predicted box in an embodiment, and (b) is the maximum matching graph of (a).
[0037] Figure 3 The results are obtained by using the OWL-ViT prediction model to detect a group of images.
[0038] Figure 4 In order to use the multi-modal large model small sample open set new energy power station inspection method of the present invention to Figure 3 The results obtained by detecting the same set of images in . DETAILED DESCRIPTION
[0039] Specific embodiments are given below in conjunction with the accompanying drawings. The specific embodiments are only used to illustrate the technical solutions of the present invention in detail and are not intended to limit the scope of protection of the present application.
[0040] The present invention provides a new energy power station inspection method based on a multimodal large model and a small sample open set (hereinafter referred to as the method, see Figure 1 ), the method comprising the following steps:
[0041] Step 1: Obtain photovoltaic inspection image sets, wind turbine inspection image sets, and tower inspection image sets from the inspection of new energy power stations. Each image set contains corresponding defect images and non-defect images. The number of images in each image set is no less than 100. Each image data in each image set includes an image with a corresponding defect category label and a corresponding text description. The defect images in each image set cover all currently known defect categories of the corresponding subject. The defect categories of the defect images in the photovoltaic inspection image set include images of hot spot defects, cracking defects, diode failure defects, contamination defects, and occlusion defects of photovoltaic panels. The defect categories of the defect images in the wind turbine inspection image set include images of dirt adhesion, oil leakage, scratches, paint peeling, rust, fading, corrosion, deformation, cracks, and wear. The defect categories of the defect images in the tower inspection image set include images of screw detachment, rust, deformation, cracks, and damaged metal accessories. The defect images and non-defect images are at least one of infrared images and visible light images.
[0042] Step 2: Build a small sample open set detection model based on a multimodal large model
[0043] This detection model includes an end-to-end ViT model (Vision Transformer), a text processing module, and an unknown class pseudo-sample mining module. The end-to-end ViT model consists of three components: an image feature embedding module (patches), a Transformer encoder, and an MLP classification head. The end-to-end ViT model serves as the base detector.
[0044] The image feature embedding module first divides the input image P0 into multiple small blocks of equal size (called patches), then adds a position code corresponding to its position in the image P0 to each small block, and then flattens each small block into a one-dimensional vector. These flattened small blocks are then embedded into a high-dimensional feature vector through a convolutional layer or linear transformation. Finally, a learnable category token is added to the front of the high-dimensional feature vector as a representative of global information, which is finally used for classification.
[0045] The Transformer encoder contains multiple encoder layers. The output of the previous encoder layer serves as the input of the next encoder layer. Each encoder layer includes a first normalization layer (Norm), a multi-head attention layer (Multi-HeadAttention), a second normalization layer (Norm), and a feedforward neural network (Feed Forward Network). The feature T0 input to an encoder layer is first processed by the first normalization layer (Norm), and the result is input to the multi-head attention layer (Multi-Head Attention). The output of the multi-head attention layer (Multi-Head Attention) is residually connected with the feature T0. The obtained feature T1 is processed by the second normalization layer (Norm) and then input into the feedforward neural network. The output of the feedforward neural network is residually connected with the feature T1, and the obtained feature T2 is used as the output of the encoder layer.
[0046] The output category token of the last encoder layer of the Transformer encoder is used as the input of the MLP classification head. The MLP classification head classifies the input and obtains the final prediction result.
[0047] The text processing module includes a text encoder (CLIP Text Encoder) from a pre-trained CLIP model and a multi-text science-based embedded query library. First, the single-text prompts (including a text description of an unknown class) for all possible results of a new energy power station inspection are modified into multi-text learnable prompts. For example, a single-text prompt such as "a picture of a wind turbine with an oil leak" is modified into multi-text learnable prompts such as "wind turbine," "oil leak," and "a picture of a wind turbine." The text encoder then extracts the feature vector of each multi-text learnable prompt. Finally, each learnable prompt and its corresponding feature vector are embedded into a data item for storage, resulting in the multi-text science-based embedded query library.
[0048] Image P0 is processed by the image feature embedding module and the Transformer encoder in turn to obtain the feature T p , feature T p The image feature f is obtained through linear projection. The category similarity between the image feature f and the multi-text science embedding query library is calculated, and then a maximum pooling operation is performed to obtain the category prediction similarity score S. The dimension of S is B × K, where B represents the number of prediction boxes and K is the total number of predicted categories, which is the sum of all known categories and the number of unknown categories.
[0049]
[0050] sin(,) represents cosine similarity, τ represents temperature coefficient, which is a set non-zero constant, set to 0.05, and image feature f is the feature obtained by linear projection of the output feature of the Transformer encoder. Multi-text science department embedding query library It is extracted by the text encoder. i The feature vector representing the i-th learnable hint in the multi-text science department embedding query base. Z is the number of learnable hints in the multi-text science department embedding query base.
[0051] The larger the element value in S, the more similar the corresponding prediction box (prediction bounding box) is to the corresponding learnable hint, that is, the learnable hint is the category label of the prediction box. According to S, the classification label of each prediction box of the image feature f is obtained.
[0052] The feature T output by the Transformer encoder p , is input to the MLP classification head, which outputs a set of N fixed-size prediction boxes and the probability that each prediction box belongs to each class label. The N prediction boxes are then input to the unknown class pseudo-sample mining module. The unknown class pseudo-sample mining module first constructs a cost matrix containing virtual unknown classes by adding unknown class virtual nodes to a bipartite graph, and then uses a bipartite graph matching algorithm to mine unknown class samples from the predicted bounding boxes.
[0053] Due to the lack of real unknown class samples during the training process, the present invention constructs a cost matrix from the predicted bounding box by adding bipartite graph unknown class virtual nodes, uses the bipartite graph matching algorithm to mine low-cost unknown class pseudo samples, optimizes the cost-aware unknown class loss function, constructs the unknown class decision boundary, and then realizes generalized small sample open set target detection. In graph theory, a matching is a set of edges, where any two edges have no common vertices. Bipartite graph matching is an important concept in graph theory, which refers to finding a way to match edges in a bipartite graph so that each vertex in the graph belongs to exactly one edge. The goal of bipartite graph matching is to find the maximum matching, that is, the graph contains the largest number of matching edges. The bipartite graph matching process of the real box and the predicted box in target detection is as follows Figure 2 As shown in (a). Figure 2 (b) is a maximum matching, which contains 3 matching edges. Matching edges are edges that do not contain common attached vertex prediction boxes, such as Figure 2As shown in (a), c(1,1) and c(2,576) are a bipartite graph match; c(1,2) and c(2,2) are not a bipartite graph match because they have a common dependent vertex prediction box 2. To mine pseudo samples of unknown classes, the present invention first adds a virtual node of unknown class - virtual box 3 - to the bipartite graph. By connecting the virtual node to each prediction box node, a cost matrix containing virtual unknown classes is constructed. Then, using a bipartite graph matching algorithm, unknown class samples are mined from the predicted bounding box. The model is trained to form the unknown class decision boundary and reject the interference of unknown classes.
[0054] In order to obtain the unknown class matching box, we first need to calculate the cost weight Q of the unknown class virtual node connected to each prediction box node match , construct a cost matrix that includes a dummy unknown class. Represents a set of N predicted labels. Assuming N is greater than the number of targets in the image, y is considered to be a set of size N filled with the true annotation label of the target and the element 0. First, a virtual label y of unknown class is used in y. u Replace an element 0, in order to find the set y and The maximum matching between N elements σ is searched at the lowest cost. σ∈G N , G N Represents the cost matrix.
[0055]
[0056] in, Represents the true label y i,i≠u and the predicted label with index σ(i) The matching cost between two nodes. The present invention uses the Kuhn-Munkres algorithm to calculate the best matching
[0057] The matching cost is equivalent to the sum of the distance between the true category prediction probability and 1 and the distance between the predicted box and the true box (true bounding box). Each element i of the true annotation label set y can be regarded as y i =(c i ,b i ), where c i is the target class label (may be ), Defined as the centroid coordinates of the ground-truth box and its height and width relative to the image size. For the predicted label indexed by σ(i), it is the class c i The probability of The prediction box is defined as L match The calculation formula is as shown in formula (3):
[0058]
[0059] in, Indicates the center distance and generalized intersection-over-union ratio between the real box and the predicted box. match Before, first define:
[0060]
[0061] Where i, j = 1, ..., N. Then the cost weight of connecting the unknown class virtual node to each prediction box node is expressed as:
[0062]
[0063] Represents the maximum cost of the connection cost between the predicted box and the real box, N gt Indicates the number of labeled objects in the image, that is, the number of predicted boxes.
[0064] In order to achieve the model in the classification process, correctly classify the known classes with hints and reject the interference of unknown classes, the unknown class pseudo samples obtained by mining are used to optimize the unknown class loss function, so that the model can distinguish between known and unknown classes in the classification stage. At the same time, the best match is considered. The intersection-over-union ratio between the unknown class candidate box and the real box and the unknown class prediction probability Therefore, we set the cost-aware unknown class loss function L ux It is defined as formula (6):
[0065]
[0066] Indicates the best match The intersection of the unknown class prediction box and the known class real box in is: Indicates the best match The predicted probability of the unknown class in c i is the target class label, are the coordinates of the center of gravity of the ground truth box and its height and width relative to the image size, For the best match The coordinates of the center of gravity of the predicted box in and its height and width relative to the image size; For the best match The index labels in The predicted label is category c i The predicted probability of the corresponding category label of each predicted box is output by the MLP classification head.
[0067] L uxFrom the perspective of optimization, the model's ability to distinguish between known and unknown classes is improved, which prompts the model to form a more compact decision boundary for unknown classes, prevents the model from mistakenly detecting unknown classes as known classes, and improves the model's open set detection performance.
[0068] Step 3: Train a small sample open set detection model based on a multimodal large model
[0069] Step 3.1 First, adjust the format of each image in the photovoltaic inspection image set, wind turbine inspection image set, and tower inspection image set in step one to meet the input format requirements of the basic detector, and then divide the defect types in each image set into three parts, the first part of which is small sample known class data, the second part is zero sample known class data, and the third part is unknown class data; small sample known class data means that the corresponding defect category on the image has been marked, and the text description of the image covers the corresponding defect category; zero sample known class data means that the corresponding defect category on the image is not marked, but the text description of the image covers the corresponding defect category; unknown class data means that the corresponding defect category on the image is marked as "unknown", and the text description of the image covers the "unknown" defect category.
[0070] As an embodiment, the photovoltaic inspection image set is set to have two categories of hot spots and cracks as small sample known categories, two categories of diode failure and pollution as zero sample known categories, and one category of occlusion as an unknown category. The tower inspection image set is set to have two categories of screw detachment and rust as small sample known categories, two categories of deformation and cracks as zero sample known categories, and one category of metal accessory damage as an unknown category. The wind turbine inspection image set is set to have three categories of dirt attachment, oil leakage, and scratches as small sample known categories, two categories of paint peeling and rust as zero sample known categories, and five categories of fading, corrosion, deformation, cracks, and wear as unknown categories.
[0071] In step 3.2, the base detector is the pre-trained ViT-B32 or ViT-L14 released by Google DeepMind. The training batch size is set to 1, the learning rate is set to 0.000003, the hyperparameters τ and ω are set to 0.05 and 1, respectively, and the maximum number of iterations is set to 100. A half-precision training strategy is used, using 16-bit floating-point numbers instead of 32-bit floating-point numbers for calculations.
[0072] Step 3.3: The image data of the photovoltaic inspection image set, the wind turbine inspection image set and the tower inspection image set are input into the pre-trained basic detector in step 3.2 respectively. An image data includes image P x and the corresponding text description W x , image P x After being processed by the image feature embedding module (patches) and the Transformer encoder, the output feature T x, feature T x After linear projection, the image feature f is obtained x .W x Input into the text processing module, first through the text encoder extraction, get the feature vector of the text description, then the text description W x And the corresponding feature vector is stored in the multi-text science embedded query library. Then calculate the image feature f x The category similarity with the multi-text science department embedding query library is subjected to the maximum pooling operation to obtain the category prediction similarity score S x , get the image feature f x The predicted label for each predicted box.
[0073] At the same time, the feature T x The input is fed into the MLP classification head, which outputs a set of N prediction boxes of a fixed size and the probability that each prediction box belongs to each prediction label. The prediction label corresponding to the maximum probability is the category label of the prediction box. The N prediction boxes are fed into the unknown class pseudo sample mining module. The unknown class pseudo sample mining module first constructs a cost matrix containing virtual unknown classes by adding unknown class virtual nodes to the bipartite graph, then uses the bipartite graph matching algorithm to mine unknown class samples from the predicted bounding boxes and calculates the cost-aware unknown class loss function L. ux The value of the unknown class loss function L ux The value of the gradient backpropagation method is used to embed the text description W in the query library of the multi-text science department x The value of the corresponding eigenvector is corrected;
[0074] According to the image P x The marked target box and the marked defect category label, category prediction similarity score S x , the MLP classification head outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each prediction label and the unknown class loss function L ux , calculate the image P x The training loss L is:
[0075]
[0076] Among them, L ce (c i ) represents the best match in the unknown class pseudo sample mining module The cross entropy classification loss is , where the best match The true category label is the label of the input image itself, and its predicted label is the similarity score S predicted according to the category x Income. represents the sum of the IoU loss and L1 regression loss between the defect location prediction box output by the MLP classification head and the true box, λ iou and λ L1 is the weighting coefficient, set to 1. ω is the unknown class loss function L ux The weighting coefficient of .
[0077] According to the image P x The training loss value is set to minimize the training loss. The AdamW optimizer is selected with the decay weight set to 0.1. The network parameters in the basic detector are updated once to complete the training of one image data. The next image data is then input into the basic detector. The network parameters of the basic detector when the previous image data is completed are used as the initial parameters for the training of the next image data. The training process for one image data is repeated continuously until the last image data is trained, completing one round of training. The network parameters of the basic detector that completed the previous round of training are used as the initial parameters for the next round of training. The training process is repeated continuously until the training round reaches the preset value, and the basic detector with the optimal parameters is obtained.
[0078] Step 4: New Energy Power Station Inspection
[0079] Photographs of photovoltaic, wind turbine, or tower inspections taken at a new energy power station are obtained. These are first formatted to conform to the input format requirements of the basic detector (resolution 512×512). These are then input into the basic detector with the optimal parameters from step three. The MLP classification head outputs the detection results, determining whether the corresponding inspected component contains defects and the defect category.
[0080] In order to verify the specific detection effect of the technical solution of the present invention, the ViT-based method OWL-ViT (Minderer M, Gritsenko A, Stone A, et al. Simple open-vocabulary object detection with vision transformers [A]. Proceedings of the European Conference on Computer Vision [C]. Tel Aviv, Israel: Springer, 2022: 728-755) was selected for comparison. In order to ensure the fairness of the experiment, the feature extraction networks of multimodal ViT-based methods such as OWL-ViT and the method of the present invention were compared using ViT-B32 and ViT-L14 respectively. Both were first trained on the same dataset and then verified on the same dataset. In the specific implementation of the present invention, three datasets were prepared: photovoltaic inspection images, wind turbine inspection images, and tower inspection images. The photovoltaic inspection images include 1,170 photovoltaic visible light images of the Guangzong new energy station in Handan and 3,520 infrared detection images of the Guangzong drone inspection, totaling 4,690 inspection images; the wind turbine inspection images mainly include 22,893 inspection images of 33 wind turbines in the Annuoji Power Station; the tower inspection images mainly include 1,198 drone tower inspection images of 31 towers in Baiyunling.
[0081] For the photovoltaic dataset, two categories, hot spots and cracks, are set as small sample known classes, two categories, diode failure and pollution, are set as zero sample known classes, and one category, occlusion, is set as an unknown class. For the tower dataset, two categories, screw detachment and rust, are set as small sample known classes, two categories, deformation and cracks, are set as zero sample known classes, and one category, metal accessory damage, is set as an unknown class. For the wind turbine dataset, three categories, dirt attachment, oil leakage, and scratches, are set as small sample known classes, two categories, paint peeling and rust, are set as zero sample known classes, and five categories, fading, corrosion, deformation, cracks, and wear, are set as unknown classes. The model input images are unified into 512×512 resolution.
[0082] For OWL-ViT, this paper pre-sets unknown class text representation and uses a pseudo sample mining algorithm based on a virtual node bipartite graph to mine unknown class pseudo samples, and trains an unknown class loss function based on logarithmic softmax, thereby achieving generalized small sample open set object detection. The evaluation indicator is mAP. fs (small sample), mAP zs (zero-sample), AR U The experimental results of WI and AOSE on the photovoltaic inspection dataset and the tower inspection dataset are shown in Table 1.
[0083] Table 1 Small sample open set evaluation results on photovoltaic inspection dataset and tower inspection dataset
[0084]
[0085] As can be seen from Table 1, the proposed method has achieved significant performance improvement in unknown class detection. In the 10-shot setting of the tower inspection dataset, the average recall rate of unknown classes of the proposed method (ViT-L14) is AR U It reached 55.1%, which is 12.8% higher than OWL-ViT. The experimental results show that the method of the present invention has strong generalization performance when facing unknown categories, and can effectively detect and identify category hints that have not been seen during the training process. In addition, the detection performance of the small sample known class and the zero sample known class of the method of the present invention showed a significant improvement. Compared with OWL-ViT, the average small sample detection result mAP of the method of the present invention is fs And zero-shot detection results mAP zs This shows that the proposed method not only shows good generalization performance for unknown classes under limited training samples, but also achieves significant performance improvement in the recognition of small-sample known classes and zero-sample known classes.
[0086] The effectiveness of our method was evaluated using a more challenging wind turbine inspection dataset, with three classes set as small-sample known classes, two classes set as zero-sample known classes, and five wind classes set as unknown classes. Compared to the two small-sample classes, two zero-sample classes, and two unknown classes used in the photovoltaic and tower inspection datasets, the wind turbine inspection dataset presents a greater challenge. The experimental results for 1-shot, 5-shot, 10-shot, and 30-shot models are shown in Table 2.
[0087] Table 2 Small sample open set evaluation results on the fan inspection dataset
[0088]
[0089] As can be seen from Table 2, the small sample performance, zero sample performance and unknown class detection performance of the method of the present invention have all been greatly improved. In the case of 10-shot small samples, the method of the present invention uses ViT-B32 as the feature extraction network, and compared with OWL-ViT, the small sample detection performance mAP is improved by 8.8%. fs , 2.0% zero-shot detection performance mAP zs , 6.6% unknown class detection performance AR U .
[0090] The visualization results on the 10-shot wind turbine inspection dataset are as follows: Figure 3 、 Figure 4As shown in Figure 2, compared with the OWL-ViT method, the method of the present invention has certain advantages in detecting unknown objects. Figure 4 As shown in the figure, in the left image, our method can detect a mark on the fan surface, while the OWL-ViT method treats it as background. In the middle image, our method detects two scratch defects. In the right image, the OWL-ViT method mistakenly detects the fan component as a scratch class, while our method detects it as an unknown class, preventing class confusion. The overall visualization analysis results show that our method has good open set class detection capabilities.
[0091] Any matters not described in the present invention are applicable to the prior art.
Claims
1. A new energy power station inspection method based on multimodal large model and small sample open set, characterized by: The method comprises the following steps: Step 1: Obtain photovoltaic inspection image sets, wind turbine inspection image sets, and tower inspection image sets from the inspection of new energy power stations. Each image set contains corresponding defect images and non-defect images. The number of images in each image set is no less than 100. Each image data in each image set includes an image with a corresponding defect category label and a corresponding text description. The defect images in each image set cover all defect categories currently known to the corresponding subject. Step 2: Build a small sample open set detection model based on a multimodal large model The detection model includes an end-to-end ViT model, a text processing module, and an unknown class pseudo sample mining module. The end-to-end ViT model consists of three parts: an image feature embedding module, a Transformer encoder, and an MLP classification head. The end-to-end ViT model is the basic detector. The image feature embedding module first divides the input image P0 into multiple small blocks of equal size, then adds a position code corresponding to its position in the image P0 to each small block, and then flattens each small block into a one-dimensional vector. These flattened small blocks are then embedded into a high-dimensional feature vector through a convolutional layer or linear transformation. Finally, a learnable category token is added to the front of the high-dimensional feature vector as a representative of global information, which is finally used for classification. The Transformer encoder contains multiple encoder layers. The output of the previous encoder layer serves as the input of the next encoder layer. Each encoder layer includes a first normalization layer, a multi-head attention layer, a second normalization layer, and a feedforward neural network. The feature T0 input to an encoder layer is first processed by the first normalization layer, and the result is input to the multi-head attention layer. The output of the multi-head attention layer is residually connected with the feature T0. The obtained feature T1 is processed by the second normalization layer and then input into the feedforward neural network. The output of the feedforward neural network is residually connected with the feature T1, and the obtained feature T2 is used as the output of the encoder layer. The output category token of the last encoder layer of the Transformer encoder is used as the input of the MLP classification head. The MLP classification head classifies the input and obtains the final prediction result. The text processing module includes a text encoder in the pre-trained CLIP model and a multi-text science department embedding query library. First, the single text prompts of all possible results of the new energy power station inspection are modified into multi-text learnable prompts. Then, the text encoder is used to extract the feature vector of each learnable prompt. Finally, each learnable prompt and the corresponding feature vector are embedded into a piece of data for storage, thereby obtaining the multi-text science department embedding query library. Image P0 is processed by the image feature embedding module and the Transformer encoder in turn to obtain the feature T p , feature T p The image feature f is obtained through linear projection. The category similarity between the image feature f and the multi-text science embedding query library is calculated, and then a maximum pooling operation is performed to obtain the category prediction similarity score S. The dimension of S is B×K, where B represents the number of prediction boxes and K is the total number of prediction categories set, which is the sum of the number of all known categories and the number of set unknown categories. sin(,) represents cosine similarity, τ represents temperature coefficient, which is a set non-zero constant, q i represents the feature vector of the i-th learnable hint in the multi-text science department embedding query base, and Z is the number of learnable hints in the multi-text science department embedding query base; The larger the element value in S is, the more similar the corresponding prediction box is to the corresponding learnable hint, that is, the learnable hint is the predicted label of the prediction box. According to S, the predicted label of each prediction box of the image feature f is obtained; The feature T output by the Transformer encoder p , is input to the MLP classification head, which outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each class label; the N prediction boxes are input to the unknown class pseudo sample mining module; the unknown class pseudo sample mining module first constructs a cost matrix containing virtual unknown classes by adding unknown class virtual nodes to the bipartite graph, and then uses the Kuhn-Munkres algorithm to search for the best match of N elements from the cost matrix with the lowest cost Then calculate the cost-aware unknown class loss, the cost-aware unknown class loss function L ux It is defined as formula (6): Indicates the best match The intersection of the unknown class prediction box and the known class real box in is: Indicates the best match The predicted probability of the unknown class in c i is the target class label, are the coordinates of the center of gravity of the ground truth box and its height and width relative to the image size, For the best match The coordinates of the center of gravity of the predicted box in and its height and width relative to the image size; For the best match The index labels in The predicted label is category c i The predicted probability of the corresponding category label of each prediction box is output by the MLP classification head; Step 3: Train a small sample open set detection model based on a multimodal large model Step 3.1 First, adjust the format of each image in the photovoltaic inspection image set, wind turbine inspection image set, and tower inspection image set in step 1 to meet the input format requirements of the basic detector. Then, divide the defect types in each image set into three parts, the first part of which is small sample known class data, the second part is zero sample known class data, and the third part is unknown class data; small sample known class data means that the corresponding defect category on the image has been labeled, and the corresponding defect category is included in the text description of the image; zero sample known class data means that the corresponding defect category on the image is not labeled, but the corresponding defect category is included in the text description of the image; unknown class data means that the corresponding defect category on the image is labeled as "unknown", and the text description of the image includes the defect category "unknown"; In step 3.2, the base detector is the pre-trained ViT-B32 or ViT-L14 released by Google DeepMind; set the training batch size to 1, set the learning rate, hyperparameters, and maximum iteration rounds; Step 3.3: The image data of the photovoltaic inspection image set, the wind turbine inspection image set and the tower inspection image set are input into the pre-trained basic detector in step 3.2 respectively. An image data includes image P x and the corresponding text description W x , image P x After being processed by the image feature embedding module and the Transformer encoder, the output feature T x , feature T x After linear projection, the image feature f is obtained x ;Describe the text W x Input into the text processing module, first through the text encoder extraction, get the feature vector of the text description, then the text description W x And the corresponding feature vector is stored in the multi-text science embedded query library; then the image feature f is calculated x The category similarity with the multi-text science department embedding query library is subjected to the maximum pooling operation to obtain the category prediction similarity score S x , get the image feature f x The predicted label for each predicted box; At the same time, the feature T x The input is to the MLP classification head, which outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each prediction label. The prediction label corresponding to the maximum probability is the category label of the prediction box; the N prediction boxes are input to the unknown class pseudo sample mining module; the unknown class pseudo sample mining module first constructs a cost matrix containing virtual unknown classes by adding virtual nodes of unknown classes in the bipartite graph, and then uses the Kuhn-Munkres algorithm to search for the best match of N elements from the cost matrix at the lowest cost, and calculates the cost-aware unknown class loss function L ux The value of the unknown class loss function L ux The value of the gradient backpropagation method is used to embed the text description W in the query library of the multi-text science department x The value of the corresponding eigenvector is corrected; According to the image P x The marked target box and the marked defect category label, category prediction similarity score S x , the MLP classification head outputs a set of N prediction boxes of fixed size and the probability that each prediction box belongs to each prediction label and the unknown class loss function L ux , calculate the image P x The training loss L is: Among them, L ce (c i ) represents the best match in the unknown class pseudo sample mining module The cross entropy classification loss is , where the best match The true category label is the label of the input image itself, and its predicted label is the similarity score S predicted according to the category x income; represents the sum of the IoU loss and L1 regression loss between the defect location prediction box output by the MLP classification head and the true box, λ iou and λ L1 is the weighting coefficient; ω is the unknown class loss function L ux The weighting coefficient of According to the image P x The training loss value is used to minimize the training loss. The AdamW optimizer is selected to update the network parameters in the basic detector once to complete the training of one image data. Then the next image data is input into the basic detector. The network parameters of the basic detector when the previous image data is completed are used as the initial parameters for the training of the next image data. The training process of one image data is repeated continuously until the last image data is trained and a round of training is completed. The network parameters of the basic detector that completed the previous round of training are used as the initial parameters for the next round of training. The training process is repeated continuously until the training round reaches the preset value and the basic detector with the optimal parameters is obtained. Step 4: New Energy Power Station Inspection Obtain inspection photos of photovoltaic, wind turbines, or towers taken at new energy power plants. First, adjust the format of the photos to meet the input format requirements of the basic detector. Then, input them into the basic detector with the optimal parameters in step 3. The MLP classification head outputs the detection results, determining whether the corresponding inspected components contain defects and the defect category.
2. The inspection method for a new energy power station based on a multimodal large model and a small sample open set according to claim 1 is characterized in that: In step 1, the defect categories of the defect images in the photovoltaic inspection image set include images of hot spot defects, cracking defects, diode failure defects, pollution defects and occlusion defects of photovoltaic panels; the defect categories of the defect images in the wind turbine inspection image set include images of dirt adhesion, oil leakage, scratches, paint peeling, rust, fading, corrosion, deformation, cracks and wear; the defect categories of the defect images in the tower inspection image set include images of screw detachment, rust, deformation, cracks and damaged metal accessories.
3. The inspection method for a new energy power station based on a multimodal large model and a small sample open set according to claim 1 is characterized in that: In step 2, the temperature coefficient τ is set to 0.
05.
4. The inspection method for a new energy power station based on a multimodal large model and a small sample open set according to claim 1 is characterized in that: Use the Kuhn-Munkres algorithm to search for the best match of N elements from the cost matrix with the lowest cost The specific process is: In order to obtain the unknown class matching box, we first need to calculate the cost weight Q of the unknown class virtual node connected to each prediction box node match , construct a cost matrix including virtual unknown classes; Represents a set of N predicted labels; assuming that N is greater than the number of targets in the image, y is regarded as a set of size N filled with the true annotation label of the target and the element 0; first, a virtual label y of unknown class is used in y u Replace an element 0, in order to find the set y and The maximum matching between N elements σ is searched at the lowest cost. σ∈G N , G N represents the cost matrix; in, Represents the true label y i,i≠u and the predicted label with index σ(i) Matching cost between two nodes; The matching cost is equivalent to the sum of the distance between the true category prediction probability and 1 and the distance between the predicted box and the true box; each element i of the true annotation label set y can be regarded as y i =(c i ,b i ), where c i is the target class label, Defined as the centroid coordinates of the ground-truth box and its height and width relative to the image size; for the predicted label indexed by σ(i), it is the category c i The probability of The prediction box is defined as L match The calculation formula is as shown in formula (3): in, Represents the center distance and generalized intersection-over-union ratio between the real box and the predicted box; in calculating Q match Before, first define: Where i, j = 1, ..., N; then the cost weight of the unknown class virtual node connected to each prediction box node is expressed as: Represents the maximum cost of the connection cost between the predicted box and the real box, N gt Indicates the number of labeled objects in the image, that is, the number of predicted boxes.
5. The inspection method for new energy power stations based on multimodal large model and small sample open set according to claim 2 is characterized in that: In step three, the photovoltaic inspection image set is set to have two categories, hot spots and cracks, as small sample known categories, diode failure and pollution, as zero sample known categories, and occlusion, as an unknown category. The tower inspection image set is set to have two categories, screw detachment and rust, as small sample known categories, deformation and cracks, as zero sample known categories, and metal accessories damage, as an unknown category. The wind turbine inspection image set is set to have three categories, dirt attachment, oil leakage, and scratches, as small sample known categories, paint peeling and rust, as zero sample known categories, and fading, corrosion, deformation, cracks, and wear, as unknown categories.
6. The inspection method for new energy power stations based on multimodal large model and small sample open set according to claim 1 is characterized in that: In step 3, the learning rate is set to 0.000003, the hyperparameters τ and ω are set to 0.05 and 1 respectively, and the maximum iteration round is set to 100.
7. The inspection method for a new energy power station based on a multimodal large model and a small sample open set according to claim 1 is characterized in that: In step 1, the defect image and the non-defect image are at least one of an infrared image and a visible light image.
Citation Information
Patent Citations
Small sample defect classification method based on retrieval enhancement and multi-modal large model
CN119622014A
Lifelong target re-identification method based on descriptive text prompt prototype compensation
CN119625285A
Multimodal data heterogeneous transformer-based asset recognition method, system, and device
US12236699B1
Performing image search based on user input using neural networks
US20230161808A1
Cited By
PID component image identification method and device, electronic equipment and storage medium
CN121527801A
Photovoltaic power station detection method based on feature decoupling domain adaptation
CN121685926A