Power transformation equipment defect identification method and device

By using a pre-built lightweight image-text matching model to perform correlation analysis on images and text descriptions of substation equipment, the problem of insufficient adaptability of existing substation equipment defect identification methods in complex environments is solved, and a more efficient defect identification effect is achieved.

CN120954017APending Publication Date: 2025-11-14FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511116893.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods for identifying defects in power equipment rely on target detection or classification models, which cannot be combined with natural language text to achieve a flexible and interpretable identification process. This results in insufficient adaptability of the models in complex and ever-changing environments and unsatisfactory identification results.

Method used

A pre-built lightweight image-text matching model is employed. This model uses a lightweight image encoder and a lightweight text encoder to perform correlation analysis between substation equipment images and predefined equipment status text descriptions, outputting substation equipment defect identification results. The model includes a lightweight image encoder to encode the image and a lightweight text encoder to encode the text description, and then uses similarity calculations to identify defects.

Benefits of technology

This invention enables the identification process of power equipment defects by combining image visual features with text semantic information, thereby improving the identification effect and adapting to complex and ever-changing real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954017A_ABST
    Figure CN120954017A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for identifying defects of power transformation equipment, which are used for solving the technical problem that the existing method for identifying the defects of the power transformation equipment limits the environment adaptability of a model, so that the identification effect is poor. The method comprises the following steps: acquiring a substation equipment image and a plurality of predefined substation equipment state text descriptions; and adopting a preset lightweight image-text matching model to carry out substation equipment defect identification according to the substation equipment image and the plurality of predefined substation equipment state text descriptions, and outputting a substation equipment defect identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power equipment technology, and in particular to a method and apparatus for identifying defects in power equipment. Background Technology

[0002] With the exponential expansion of the power system, the operational safety of key equipment such as substations and transmission lines has become a core element in ensuring the stable operation of the power grid. As important components of substations, the operating status of equipment such as high-voltage capacitors and expanders is directly related to the safety and reliability of the entire power system.

[0003] During long-term operation, electrical equipment may experience wear and tear due to various factors such as ambient temperature, internal aging, and electrical stress. Over time, this wear and tear can accumulate and pose significant safety hazards, directly leading to damage to transmission lines and power outages for users. For some critical electrical equipment, given the inherent inherent danger of electrical resources, malfunctions in these areas can pose even greater risks and have more severe consequences. For example, high-voltage capacitors may bulge, while expansion joints may experience overshooting. These apparent structural changes not only indicate potential equipment failures but can also lead to power accidents and even endanger personal safety. Therefore, the detection of abnormalities in power equipment is an urgent and crucial task.

[0004] Existing methods for identifying defects in power equipment largely rely on target detection or classification models. These models can only output fixed defect labels and are difficult to combine with natural language text to achieve a flexible and interpretable identification process. This limitation not only weakens the model's adaptability to complex and changing environments but also makes the identification results lack semantic traceability, ultimately leading to unsatisfactory identification performance in real-world scenarios. Summary of the Invention

[0005] This invention provides a method and apparatus for identifying defects in power equipment, which addresses the technical problem that existing methods for identifying defects in power equipment limit the environmental adaptability of the model, resulting in poor identification performance.

[0006] The first aspect of this invention provides a method for identifying defects in power equipment, comprising:

[0007] Acquire images of substation equipment and multiple predefined text descriptions of substation equipment status;

[0008] A pre-built lightweight image-text matching model is used to identify substation equipment defects based on the substation equipment images and multiple predefined substation equipment status text descriptions, and the substation equipment defect identification results are output.

[0009] Optionally, the pre-set lightweight image-text matching model includes a lightweight image encoder and a lightweight text encoder; the step of using the pre-set lightweight image-text matching model to identify substation equipment defects based on the substation equipment image and multiple predefined substation equipment status text descriptions, and outputting substation equipment defect identification results, includes:

[0010] The lightweight image encoder is used to encode the substation equipment image and output the image features corresponding to the substation equipment image.

[0011] The lightweight text encoder is used to encode the text descriptions of the predefined substation equipment status respectively, and outputs the text features of the text descriptions of the predefined substation equipment status.

[0012] Based on the image features and the text features, substation equipment defects are identified, and the substation equipment defect identification results are output.

[0013] Optionally, the lightweight image encoder includes a main intervention processing layer, an inverse residual module, multiple attention mechanism low-rank adaptive image modules, and an image projection head; the step of using the lightweight image encoder to encode the substation equipment image and outputting the image features corresponding to the substation equipment image includes:

[0014] The images of the substation equipment are preprocessed to output images that conform to the input dimensions of the model;

[0015] The main intervention processing layer performs preliminary downsampling on the image that conforms to the model input size to generate downsampled image features.

[0016] The downsampled image features are used as input to the inverse residual module, and the inverse residual image features are output.

[0017] Perform convolution operation on the inverted residual image features to output local image features;

[0018] Multiple attention-based low-rank adaptive image modules are used to iteratively enhance the local image features to generate global image features;

[0019] The global image features and the local image features are concatenated to output the concatenated image features, and the concatenated image features are convolved to output the target fused image features.

[0020] The target fused image features are globally pooled to output the target globally pooled image features, and vector extraction is performed on the target globally pooled image features to generate an image representation vector;

[0021] The image representation vector is projected using an image projection head to generate image features corresponding to the substation equipment image.

[0022] Optionally, the step of using the downsampled image features as input to the inverse residual module and outputting inverse residual image features includes:

[0023] The downsampled image features are channel-expanded to generate expanded image features;

[0024] Perform depthwise convolution on the extended image features to output depthwise convolution image features;

[0025] The deep convolutional image features are projected to generate projected image features;

[0026] The projected image features and the downsampled image features are added together to output the inverse residual image features.

[0027] Optionally, the low-rank adaptive image module of the attention mechanism includes a multi-head attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning; the feature processing procedure of the low-rank adaptive image module of the attention mechanism is as follows:

[0028] The input image features input to the low-rank adaptive image module of the attention mechanism are divided into blocks, and multiple image sub-features are output.

[0029] Each of the image sub-features is flattened to generate multiple image sub-feature vectors;

[0030] Add positional encoding to each of the image sub-feature vectors to output multiple image sub-feature codes;

[0031] A multi-head attention module based on low-rank matrix fine-tuning is used to output the multi-head attention image features corresponding to each image sub-feature encoding according to the image sub-feature encoding;

[0032] Each of the multi-head attention image features is used as input to a feedforward network based on low-rank matrix fine-tuning, and the corresponding feedforward image features are output.

[0033] The feedforward image features are averaged and aggregated to output the aggregated image features.

[0034] The aggregated image features are embedded and projected to generate output image features.

[0035] Optionally, the lightweight text encoder includes multiple attention-based low-rank adaptive text modules and a tagging convergence layer; the lightweight text encoder is used to encode the text of each of the predefined substation equipment status text descriptions, and outputs the text features of each of the predefined substation equipment status text descriptions, including:

[0036] Each of the predefined substation equipment status text descriptions is preprocessed to output the word sequence corresponding to each of the predefined substation equipment status text descriptions;

[0037] Each of the aforementioned word sequences is concatenated with a preset learnable prompt to output the target word sequence corresponding to each of the aforementioned word sequences;

[0038] A low-rank adaptive text module employing multiple attention mechanisms performs iterative semantic enhancement based on each target word sequence, outputting word sequence features corresponding to each target word sequence;

[0039] Each of the word sequence features is subjected to layer normalization to generate word sequence normalization features corresponding to each of the word sequence features;

[0040] The normalized features of each word sequence are input into the labeling convergence layer, and the text features corresponding to each normalized feature of the word sequence are output.

[0041] Optionally, the attention mechanism low-rank adaptive text module includes a self-attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning; the feature processing procedure of the attention mechanism low-rank adaptive text module is as follows:

[0042] The input text features input to the low-rank adaptive text module of the attention mechanism are processed by a self-attention module based on low-rank matrix fine-tuning to calculate attention and output the self-attention text features corresponding to the input text features.

[0043] A feedforward network based on low-rank matrix fine-tuning is used to perform a nonlinear transformation on the self-attention text features, and the feedforward text features are output.

[0044] The feedforward text features and the input text features are added together to output the text addition features. The text addition features are then layer-normalized to generate the output text features.

[0045] Optionally, the step of identifying substation equipment defects based on the image features and each of the text features, and outputting the substation equipment defect identification result, includes:

[0046] The similarity between the image features and each of the text features is calculated, and the similarity between the image features and each of the text features is output.

[0047] The predefined substation equipment status text description associated with the text feature corresponding to the highest similarity is selected as the substation equipment defect identification result.

[0048] Optionally, the model training process of the pre-set lightweight image-text matching model is as follows:

[0049] Acquire images of substation equipment for model training;

[0050] The images of substation equipment used for model training and the predefined text descriptions of substation equipment status are preprocessed to output image-text pairings.

[0051] An initial lightweight image-text matching model is used to pair the image and text, and the training similarity corresponding to the image-text pair is output.

[0052] Substitute the training similarity into the preset loss function and take the derivative to output the model gradient;

[0053] The model parameters of the initial lightweight image-text matching model are updated using the model gradient to determine the intermediate lightweight image-text matching model, and the number of model updates is counted in real time.

[0054] Determine whether the number of model updates has reached the preset number of training iterations;

[0055] If so, the intermediate lightweight image-text matching model is used as the trained preset lightweight image-text matching model.

[0056] A second aspect of the present invention provides a defect identification device for power equipment, comprising:

[0057] The acquisition module is used to acquire images of substation equipment and multiple predefined text descriptions of substation equipment status;

[0058] The identification module is used to identify substation equipment defects based on the substation equipment images and multiple predefined substation equipment status text descriptions using a pre-set lightweight image-text matching model, and output the substation equipment defect identification results.

[0059] As can be seen from the above technical solutions, the present invention has the following advantages:

[0060] The above-mentioned technical solution of the present invention provides a method for identifying defects in substation equipment. This method acquires images of substation equipment and multiple predefined text descriptions of substation equipment status. A pre-set lightweight image-text matching model is used to identify defects in the substation equipment based on the images and the predefined text descriptions, and the identification result is output. Based on this solution, by acquiring images of substation equipment and multiple predefined text descriptions of equipment status, and using a pre-set lightweight image-text matching model to perform correlation analysis between the two, the process of identifying and outputting results for substation equipment defects is achieved. This invention enables the identification process to utilize both the visual features of the images and the semantic information of the text, effectively improving the identification effect in real-world scenarios. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a flowchart of the steps of a method for identifying defects in power equipment according to Embodiment 1 of the present invention;

[0063] Figure 2 This is a schematic diagram of the structure of the pre-set lightweight image-text matching model provided in Embodiment 1 of the present invention;

[0064] Figure 3 This is a flowchart illustrating a method for identifying defects in power equipment according to Embodiment 1 of the present invention.

[0065] Figure 4 This is a flowchart illustrating the steps of training a pre-built lightweight image-text matching model according to Embodiment 2 of the present invention.

[0066] Figure 5 This is a flowchart of text enhancement and segmentation provided in Embodiment 2 of the present invention;

[0067] Figure 6 This is a structural block diagram of a power equipment defect identification device provided in Embodiment 3 of the present invention. Detailed Implementation

[0068] This invention provides a method and apparatus for identifying defects in power equipment, which addresses the technical problem that existing methods for identifying defects in power equipment limit the environmental adaptability of the model, resulting in poor identification performance.

[0069] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0070] Terminology Explanation:

[0071] CLIP (Contrastive Language-Image Pretraining) is a model that excels in open-world representation across domains and modalities, serving as the foundation for various visual and multimodal tasks. The model achieves visual-linguistic feature alignment: enabling visual features (such as image features) and linguistic features (such as textual descriptions) to match and correlate within the same semantic space. The CLIP model consists of an Image Encoder and a Text Encoder, which respectively acquire image and text information.

[0072] MobileCLIP model: This is a lightweight version of the CLIP model. It retains the structure and comparison calculation method of the CLIP model, but has been lightweighted to adapt to edge devices. The overall parameters of the model, as well as the structure of the image encoder and text encoder, have been adjusted to the corresponding lightweight versions.

[0073] Transformer: The purpose of this structure is to deeply understand the content of images or text, extract the most critical semantic features, process the "relationships" in images or text, learn which parts of an image are related to each other, and which words in a text are more important to the current task, and then process, understand and extract features from the input information.

[0074] Patch: Turns an image into a "sequence" so that it can be processed by the Transformer. Because images are not natural "sequences" like text, in order to use the Transformer to process images, the whole image is first cut into small pieces (Patch), just like dividing a large jigsaw puzzle into small pieces.

[0075] Embedding transforms raw image patches or text words into semantic vectors that machines can understand and compute. It converts text or image patches into numerical vectors, which are how the model understands the semantics. For image patch embedding, each image patch is typically transformed into a vector using a small convolutional or linear layer. For text word vectors, embedding maps "words" to a vector space; for example, "transformer" and "bulge" would be mapped to different but semantically similar locations.

[0076] Position embedding: Supplements "sequence" or "spatial location information" to help the model understand the overall structure. Since the Transformer itself does not know the position of each element in the sequence (such as where the 1st word, 5th word, or 20th image patch is), it is necessary to add a positional encoding to each input to tell the model where these words or image patches appear.

[0077] Please see Figure 1 , Figure 1 This is a flowchart illustrating the steps of a method for identifying defects in power equipment according to Embodiment 1 of the present invention.

[0078] The present invention provides a method for identifying defects in power equipment, comprising:

[0079] Step 101: Obtain images of substation equipment and multiple predefined text descriptions of substation equipment status.

[0080] The predefined substation equipment status text descriptions are text descriptions reconstructed and expanded using natural language by humans or rule-based systems based on core descriptions such as "capacitor bulging," "expansioner overshooting," and "equipment normal," combined with text templates. For example, for "capacitor bulging," the following text descriptions can be generated: "A capacitor bulging phenomenon has been detected on the equipment's exterior," "The current equipment status is abnormal, and capacitor bulging is suspected," and "The capacitor bulging problem is quite obvious, and it is necessary to promptly investigate potential operational safety hazards."

[0081] The images of substation equipment are obtained frame by frame from the video stream data returned by the substation monitoring camera.

[0082] Step 102: Use a pre-set lightweight image-text matching model to identify substation equipment defects based on substation equipment images and multiple predefined substation equipment status text descriptions, and output the substation equipment defect identification results.

[0083] It should be noted that you should refer to [link / reference]. Figure 2The pre-built lightweight image-text matching model mainly consists of a lightweight image encoder (MobileViT-LoRA, Mobile Vision Transformer with Low-Rank Adaptation) and a lightweight text encoder (MobileBERT-LoRA, Mobile Bidirectional Encoder Representations from Transformers with Low-Rank Adaptation). The lightweight image encoder mainly consists of a main intervention processing layer (stem), an inverted residual block, multiple low-rank adaptive image modules with attention mechanisms, and an image projection head. The low-rank adaptive image modules with attention mechanisms mainly consist of a multi-head attention module based on low-rank matrix fine-tuning and a feed-forward network based on low-rank matrix fine-tuning. The lightweight text encoder mainly consists of multiple low-rank adaptive text modules with attention mechanisms and a label pooling layer (cls pooling). The low-rank adaptive text modules with attention mechanisms mainly consist of a self-attention module based on low-rank matrix fine-tuning (Attention-Lora) and a feed-forward network based on low-rank matrix fine-tuning (FFN-Lora). It consists of (Adaptation).

[0084] It is worth mentioning that, considering the performance limitations of on-site equipment and training costs, this invention uses the MobileCLIP model as the multimodal model. This model is a lightweight version of the image-text comparison multimodal model CLIP, i.e., a pre-built lightweight image-text matching model. This invention also considers the relatively small amount of defective data and optimizes and fine-tunes MobileCLIP to adapt it to the current task. This invention adapts the visual model and language model (i.e., the lightweight image encoder and lightweight text encoder) within MobileCLIP separately. Furthermore, this invention fine-tunes MobileViT to adapt it to few-sample training tasks.

[0085] Furthermore, step 102 may include the following sub-steps S21-S23:

[0086] Step S21: Use a lightweight image encoder to encode the substation equipment image and output the image features corresponding to the substation equipment image.

[0087] Specifically, step S21 may include the following sub-steps S211-S218:

[0088] Step S211: Preprocess the substation equipment image and output an image that conforms to the input size of the model.

[0089] It should be noted that the original image input into the model, i.e., the substation equipment image, is... First, it needs to be preprocessed to obtain an image that conforms to the model input size, i.e., an image that conforms to the model input size. :

[0090] ;

[0091] Where H and W are the height and width of the substation equipment image; preprocess is the preprocessing operation that reshapes the input images of different sizes into a specific 224×224 shape to facilitate model processing.

[0092] Step S212: Perform preliminary downsampling on the image that conforms to the model input size through the main intervention processing layer to generate downsampled image features.

[0093] It should be noted that the obtained image x, which conforms to the model input size, is input into the STEM layer (the main intervention processing layer) for low-level feature processing. The STEM layer mainly performs preliminary downsampling on the input image, extracting low-level features such as edges and textures to obtain downsampled image features. This process can be represented as:

[0094] ;

[0095] in, This represents a 3×3 convolutional layer; BN is the normalization function; SiLU is the activation function.

[0096] Step S213: Use the downsampled image features as input to the inverse residual module and output the inverse residual image features.

[0097] Further, step S213 may include the following sub-steps S2131-S2134:

[0098] Step S2131: Perform channel expansion on the downsampled image features to generate expanded image features;

[0099] Step S2132: Perform depthwise convolution on the extended image features and output the depthwise convolution image features;

[0100] Step S2133: Perform a projection operation on the depth convolution image features to generate projected image features;

[0101] Step S2134: Add the features of the projected image and the downsampled image to output the inverse residual image features.

[0102] It should be noted that the obtained features (downsampled image features) are input into the Inverted ResidualBlock (IRB) module, i.e., the inverse residual module, to model local contextual information and perform nonlinear feature transformation using a lightweight structure. This process extracts more abstract semantics while preserving the input information. Specifically, the downsampled image features are first channel-expanded to obtain expanded image features. :

[0103] ;

[0104] Then, depthwise convolution is performed to obtain the depthwise convolution image features. :

[0105] ;

[0106] Then perform a projection operation to obtain the features of the projected image. :

[0107] ;

[0108] Finally, residual connections are performed to obtain the inverse residual image features. :

[0109] ;

[0110] Where exp is expansion; project is projection; and DWConv is depthwise convolution.

[0111] Step S214: Perform convolution operation on the inverted residual image features to output local image features.

[0112] It should be noted that the obtained inverse residual image features The input is fed into the MobileViT Block (MobileVision Transformer Block) to fuse local features and global contextual information. Simultaneously, within this structure, the model's ability to model long-range dependencies is enhanced. First, a 3×3 convolution is used to obtain information about the local receptive field, i.e., local image features. : .

[0113] Step S215: Use multiple attention mechanism low-rank adaptive image modules to perform iterative image enhancement on local image features to generate global image features.

[0114] Optionally, the feature processing procedure of the low-rank adaptive image module of the attention mechanism is as follows:

[0115] The input image features to the low-rank adaptive image module of the attention mechanism are divided into blocks, and multiple image sub-features are output.

[0116] Each image sub-feature is flattened to generate multiple image sub-feature vectors;

[0117] Add positional encoding to each image sub-feature vector to output multiple image sub-feature codes;

[0118] A multi-head attention module based on low-rank matrix fine-tuning is used to output the multi-head attention image features corresponding to each image sub-feature encoding according to the sub-feature encoding of each image.

[0119] Each multi-head attention image feature is used as input to a feedforward network based on low-rank matrix fine-tuning, and the output is the feedforward image feature corresponding to each multi-head attention image feature.

[0120] The features of each feedforward image are averaged and aggregated to output the aggregated image features.

[0121] Embedded projection is performed on the aggregated image features to generate output image features.

[0122] The input image features are feature maps input to the low-rank adaptive image module of the attention mechanism. It can be understood that the input image features can correspond to any feature map input to the low-rank adaptive image module of the attention mechanism for image processing during the training or detection of the pre-set lightweight image-text matching model.

[0123] Image sub-features, image sub-feature vectors, image sub-feature encoding, multi-head attention image features, feedforward image features, and aggregated image features are all intermediate images generated in the low-rank adaptive image module of the attention mechanism.

[0124] The output image features are feature maps output by the low-rank adaptive image module of the attention mechanism. It can be understood that they can correspond to any output feature map of the low-rank adaptive image module of the attention mechanism after image processing during the training or detection of the pre-set lightweight image-text matching model.

[0125] It should be noted that multiple attention mechanism low-rank adaptive image modules are used to iteratively enhance the semantics of local image features to generate global image features; where the output of the Nth attention mechanism low-rank adaptive image module is the input of the N+1th attention mechanism low-rank adaptive image module, and the global image features are the output of the last attention mechanism low-rank adaptive image module.

[0126] Furthermore, for the feature processing of the low-rank adaptive image module of the attention mechanism, the input feature map is first converted into an input patch sequence for the Transformer module (a multi-head attention module based on low-rank matrix fine-tuning), with a patch size of p×p: Where N is the number of patches and C is the number of channels; The image is a sub-feature vector; Unfold is the operation of dividing the feature map into individual patches and flattening them into vectors.

[0127] Furthermore, the obtained patch sequence (i.e., multiple image sub-feature vectors) is input into the Transformer-LoRA module for Positional Encoding, adding positional encoding to each patch and performing the following calculations:

[0128] ;

[0129] in, It is a linear matrix. D represents the number of output channels, and C represents the number of input channels; For position encoding, N is the number of patches; Each patch represents a feature, i.e., an image sub-feature vector; This refers to the encoded features, i.e., the image sub-feature encoding.

[0130] Furthermore, this invention introduces a LoRA module to fine-tune the weight matrices of the multi-head attention module and the feedforward network, resulting in a multi-head attention module and a feedforward network based on low-rank matrix fine-tuning. LoRA aims to incrementally adjust the weight matrix W through low-rank matrix decomposition, without directly modifying the original weights, but by adding a low-rank matrix product term. ,in, The adjusted weights are truly heartbreaking. As an increment, A and B are two newly added low-rank matrices. During training, only their corresponding contents need to be updated, while the weight matrix W remains frozen.

[0131] Furthermore, the image sub-feature encoding is processed through a multi-head attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning. The output of the k-th layer is then fine-tuned by adding LoRA. for: ;in, , It is the output of the k-th layer after fine-tuning with LoRA, that is, the feedforward text features of the k-th layer. Output for layer k-1 The coarse features obtained after inputting into layer k need to be processed by a feedforward neural network to obtain the output of layer k. ; To improve the output of the (k-1)th layer after fine-tuning LoRA; FFN represents the feedforward network, and x represents the input of the feedforward network. , Let be the weight matrix of the linear transformation in FFN. , This is a low-rank matrix related to LoRA (Low-Rank Adaptation). , This represents the bias term for the linear transformation in FFN; MSA stands for Multi-Head Attention Module, i.e., the operation of the multi-head attention mechanism, and its calculation formula is:

[0132] ;

[0133] Where Q, K, and V represent the query, key, and value feature matrices respectively during the attention calculation process; is the dimension of the key feature matrix; As a dynamic information filter, it extracts relevant features from values ​​(V) and performs weighted fusion by calculating the matching degree between the query (Q) and the key (K); after fine-tuning with the addition of the LoRA module, each header h has:

[0134] , , ;

[0135] in, The matrix is ​​used to calculate the corresponding Q, K, and V matrices. ; This is a low-rank matrix used to calculate the corresponding Q, K, and V matrices. x is the input to MSA.

[0136] Then, multiple headers are concatenated and mapped:

[0137] ;

[0138] in, is the output of the h-th attention head; x is the input of MSA; Concat is for concatenation; This is the weight matrix used for linear mapping of the spliced ​​multi-head output.

[0139] Furthermore, the output of the feedforward network is subjected to Global Pooling and Embedding Projection to generate the final encoding:

[0140] ;

[0141] ;

[0142] in, Aggregate image features; This represents the number of patches. Let be the feature vector of the i-th image patch (image sub-feature vector) in the L-th layer, i.e., the feedforward image feature; D is... The dimension; To output image features; This is a LayNorm layer used for layer normalization; Encode the dimensions of the Transformer's output. d is The dimension of a vector.

[0143] Step S216: Concatenate the global image features and local image features to output the concatenated image features, and perform convolution operation on the concatenated image features to output the target fused image features.

[0144] Step S217: Perform global pooling on the target fused image features, output the target global pooled image features, and extract vectors from the target global pooled image features to generate image representation vectors.

[0145] Step S218: Project the image representation vector using an image projection head to generate image features corresponding to the substation equipment image.

[0146] It should be noted that the output image features of the last attention mechanism's low-rank adaptive image module are used as global image features. ,Right now fold indicates a pair The process involves restoring the spatial structure, followed by fusing local and global information. This involves fusing the output features of the CNN and Transformer, specifically concatenating global and local image features to output a concatenated image feature. This convolution operation is then performed on the concatenated image feature to output the target fused image feature. :

[0147] ;

[0148] Furthermore, Global Pooling is performed to extract the image representation vector from the feature map, generating the image semantic vector, i.e., the image representation vector:

[0149] ;

[0150] Finally, the image vectors are shared into a multimodal shared space after passing through the image projection head, resulting in image features corresponding to the substation equipment images:

[0151] ;

[0152] in, Let be the image representation vector of the i-th substation equipment image, representing the global pooling result; H represents the unnormalized image features; H is the height of the feature map; W is the width of the feature map. This is the feature vector of the output feature map at a position with height h and width w; For fully connected layer coefficients; To calculate the bias; The image features are the corresponding to the image of the i-th substation equipment; at this point, the image encoder obtains the output and maps it to the multimodal space.

[0153] Step S22: Use a lightweight text encoder to encode the text descriptions of the status of each predefined substation equipment, and output the text features of each predefined substation equipment status text description.

[0154] It should be noted that when using a lightweight text encoder (MobileViT-LoRA, a variant of text encoding) to encode the predefined text descriptions of substation equipment status, the text is first split into discrete word sequences, converted into initial vectors through an embedding layer, and then positional encoding is superimposed to preserve word order information. Next, a lightweight Transformer module (combining depthwise separable convolution and attention mechanisms) is used for contextual semantic modeling. Simultaneously, LoRA technology is utilized to enhance feature representation capabilities while maintaining the model's lightweight nature, ultimately outputting fixed-dimensional text feature vectors to achieve a compact semantic representation of equipment status descriptions. This significantly reduces the number of model parameters and computational complexity while ensuring encoding accuracy, adapting to the resource constraints of substation edge computing scenarios. Contextual modeling improves the semantic discriminativeness of text features, accurately capturing subtle differences between different equipment status descriptions. The LoRA mechanism enhances the model's adaptability to professional domain texts, improving the targeting and robustness of feature encoding.

[0155] Specifically, step S22 may include the following sub-steps S221-S225:

[0156] Step S221: Preprocess the text descriptions of the status of each predefined substation equipment and output the word sequence corresponding to each predefined text description of the status of substation equipment.

[0157] It should be noted that the input sentence or text, i.e., the predefined text description of substation equipment status, must first be processed into a token format that the text encoder can read. For the input sentence T = "capacitor bulge", it will first pass through the predefined word segmentation unit of the text encoder to extract individual characters, map words to word vectors, and add position vectors to them. That is, each word is converted into a vector according to the existing mapping table and a position code is added, so that the model can recognize the order of words in the sentence and obtain the word sequence corresponding to the predefined substation equipment status text description. : ,in, For position encoding, Represents a mapping, It is a feature or representation related to the i-th position in the original input, where i is the position index in the word sequence.

[0158] Step S222: Concatenate each word sequence with the preset learnable prompts to output the target word sequence corresponding to each word sequence.

[0159] It should be noted that, since the focus is on improving performance with a small number of samples, a prompt tuning part is first added to the text encoder. This involves adding a trainable embedding sequence before each word vector, directly concatenating it to the input token. This guides the model to focus on the task; essentially, it inserts a learnable prompt (a pre-set learnable prompt), represented as... And insert it before the input sequence: ,in, The target word sequence corresponding to the word sequence; This is the k-th element in the pre-set learnable hints.

[0160] Step S223: Employ multiple attention mechanisms and low-rank adaptive text modules to perform iterative semantic enhancement based on each target word sequence, and output the word sequence features corresponding to each target word sequence.

[0161] It should be noted that multiple attention-based low-rank adaptive text modules are used to perform iterative semantic enhancement on each target word sequence. Each module captures semantic dependencies within the word sequence using an attention mechanism, while simultaneously employing low-rank adaptive technology (LoRA) to efficiently update module parameters, enhancing semantic representation capabilities without significantly increasing computational overhead. During the iteration process, the features output by the previous module serve as input to the next module, progressively improving the semantic richness and expressive power of the word sequence features, ultimately yielding the word sequence features corresponding to each target word sequence. Specifically, each target word sequence undergoes iterative semantic enhancement processing, and the output text features of the final attention-based low-rank adaptive text module are the corresponding word sequence features for the target word sequence. Through multi-module iterative processing, the semantic information of the word sequences is effectively enhanced, enabling features to more accurately reflect the semantic connotation of the text. The low-rank adaptive technology improves model performance while avoiding parameter explosion, ensuring computational efficiency. The synergistic effect of multiple modules enhances the model's adaptability to complex semantic scenarios and improves the usability of word sequence features in subsequent tasks.

[0162] Optionally, the feature processing procedure for the low-rank adaptive text module of the attention mechanism is as follows:

[0163] The input text features of the low-rank adaptive text module of the input attention mechanism are processed by the self-attention module based on low-rank matrix fine-tuning, and the self-attention text features corresponding to the input text features are output.

[0164] A feedforward network based on low-rank matrix fine-tuning is used to perform nonlinear transformation on self-attention text features and output feedforward text features.

[0165] The feedforward text features and the input text features are added together to output the added text features. Then, the added text features are layer normalized to generate the output text features.

[0166] The input text features are feature maps input to the low-rank adaptive text module of the attention mechanism. It can be understood that the input text features can correspond to any feature map input to the low-rank adaptive text module of the attention mechanism for image processing during the training or detection of the pre-set lightweight image-text matching model.

[0167] Self-attention text features, feedforward text features, and text summation features are all intermediate graphs generated in the low-rank adaptive text module of the attention mechanism.

[0168] The output text features are feature maps output by the low-rank adaptive text module of the attention mechanism. It can be understood that they can correspond to any output feature map of the low-rank adaptive text module of the attention mechanism after image processing during the training or detection of the pre-set lightweight image-text matching model.

[0169] The self-attention module based on low-rank matrix fine-tuning is

[0170] It should be noted that each target word sequence will pass through multiple MobileBERT-LoRA Blocks to obtain corresponding output features, i.e., word sequence features. The Block used here is the MobileBERT-LoRA fine-tuned for the text encoder in this invention for the case of small samples. The principle is still to add a LoRA module to adapt to the case of a small number of samples. That is, the LoRA module of this invention fine-tunes the weight matrix of the self-attention module and the feedforward network to obtain a self-attention module and a feedforward network fine-tuned based on a low-rank matrix. The specific fine-tuning part and the overall workflow are as follows:

[0171] MobileBERT-LoRA first performs self-attention calculations, allowing each word to interact with other words to obtain global semantic information:

[0172] , , ;

[0173] in, For the input of the self-attention module based on low-rank matrix fine-tuning, represents the input text features; , , , This represents the weight matrix before adding the LoRA module. To add the weight matrix after fine-tuning the LoRA module, The two newly added low-rank matrices, Q, K, and V, represent the query, key, and value feature matrices respectively during the attention calculation process. After calculating the attention, the output of the self-attention module, which is fine-tuned based on a low-rank matrix, is input into the feedforward network, which is also fine-tuned based on a low-rank matrix, to perform non-linear changes on the position of each word, thereby improving expressive power. z is a sequence The resulting sequence after processing through an FFN layer is the feedforward text feature, where ReLU is the activation function. For self-attention text features, represents the output of the self-attention module based on low-rank matrix fine-tuning. The weight matrix is ​​used for linear transformation in a feedforward network based on low-rank matrix fine-tuning. To optimize the linear matrix, the original linear matrix is ​​W1. To improve performance with a smaller dataset, the LoRA module is used for fine-tuning, calculated as follows:

[0174] ;

[0175] in, Here are the weight parameters for the LoRA module; These are the two matrices used for fine-tuning corresponding to the LoRA module.

[0176] It is worth mentioning that this invention only requires LoRA fine-tuning of the inner layer, while the outer layer matrix does not need to be fine-tuned. This significantly reduces the number of parameters and computational resource consumption during the fine-tuning process, enabling the model to efficiently adapt in resource-constrained edge computing environments (such as local terminals of substation monitoring systems). Fixing the outer layer matrix effectively preserves the general knowledge base of the pre-trained model, ensuring that the performance improvement of the model on the target task after fine-tuning does not come at the expense of basic capabilities. The low-rank characteristic of LoRA enhances the model's adaptability to tasks, improving the accuracy of semantic understanding of equipment status text while reducing the risk of overfitting, giving the model stronger generalization ability in professional fields such as substation equipment status monitoring.

[0177] Further, residual joins and layer normalization operations are then performed:

[0178] ;

[0179] in, The output text features are defined by LayerNorm, which represents the normalization operation; Sublayer represents the subsequent multiple layers of MobileBERT-LoRA Block modules, i.e., the low-rank adaptive text module of the attention mechanism.

[0180] Step S224: Perform layer normalization on each word sequence feature to generate word sequence normalized features corresponding to each word sequence feature.

[0181] Step S225: Input the normalized features of each word sequence into the labeling convergence layer, and output the text features corresponding to the normalized features of each word sequence.

[0182] It should be noted that the output of the Nth attention mechanism low-rank adaptive text module... As the N+1th attention mechanism, the low-rank adaptive text module The output of the last attention mechanism, the low-rank adaptive text module, is... The word sequence features corresponding to the target word sequence are then used to obtain the text embedding representing the entire sentence through layer normalization and the cls pooling module (labeled pooling layer), which is the text feature corresponding to the word sequence normalization feature:

[0183] ;

[0184] This formula means that for the output sequence, i.e., the output of the last attention mechanism low-rank adaptive text module, ... The entire input is pooled using the cls pooling module, but only the vector corresponding to the cls position in the entire sequence is taken as the output vector, i.e., the text feature. cls is the token sequence initially added to the model, and is used during token extraction and construction. At that time, the model will automatically initialize a trainable variable cls and add it to the model. Since subsequent operations are based on a self-attention mechanism, the initialized `cls` variable will learn information from the entire sentence and represent the whole sentence information. Therefore, only the vector corresponding to `cls` is taken as the output here. This represents retrieving a vector from a specific position within a given sentence. `cls` is a variable containing all the information from that sentence. The final output of the text encoder is... .

[0185] Step S23: Based on image features and text features, identify defects in the power equipment and output the defect identification results.

[0186] It should be noted that after the image and text are encoded by the image encoder and text encoder respectively to obtain the image-text multimodal feature vectors, the matching relationship between the image and text information is learned, and the feature vectors obtained by the image encoder and the feature vectors obtained by the text encoder are matched.

[0187] Specifically, step S23 may include the following sub-steps S31-S32:

[0188] Step S31: Calculate the similarity between the image features and each text feature, and output the similarity between the image features and each text feature;

[0189] Step S32: Select the predefined substation equipment status text description associated with the text feature corresponding to the highest similarity as the substation equipment defect identification result.

[0190] It should be noted that the formula for calculating similarity is as follows:

[0191] ;

[0192] in, This represents the similarity between image features and text features.

[0193] Furthermore, based on the output text information (predefined text descriptions of substation equipment status), it is determined whether the output is normal or abnormal. Since this patent mainly targets the identification of two important substation equipment, capacitors and expanders, the output content is relatively fixed. Therefore, the category information to be identified can be given in advance, and the categories of text that belong to abnormal output can be manually pre-classified.

[0194] For comparison of technical effects, existing technologies can be used as a reference. With social development and the country's large-scale investment in power infrastructure, electricity has become an essential energy source for everyone every day. At the same time, power equipment has been developed and popularized in every corner of the country. Whether in core cities or remote villages, the power system plays an important role. Because of the continuous expansion of the power system, the safe operation of key equipment such as substations and transmission lines has become the core link to ensure the stable operation of the power grid. As an important component of substations, the operating status of high-voltage capacitors and expanders is directly related to the safety and reliability of the entire power system.

[0195] During long-term operation, electrical equipment may experience wear and tear due to various factors such as ambient temperature, internal aging, and electrical stress. Over time, this wear and tear can accumulate and pose significant safety hazards, directly leading to damage to transmission lines and power outages for users. For some critical electrical equipment, given the inherent inherent danger of electrical resources, malfunctions in these areas can pose even greater risks and have more severe consequences. For example, high-voltage capacitors may bulge, while expansion joints may experience overshooting. These apparent structural changes not only indicate potential equipment failures but can also lead to power accidents and even endanger personal safety. Therefore, the detection of abnormalities in power equipment is an urgent and crucial task.

[0196] Currently, power companies primarily assess equipment condition through manual inspections and infrared imaging. However, manual inspections suffer from low efficiency, high error rates, and the inability to achieve 24 / 7 intelligent monitoring. Furthermore, manual inspections of remote power equipment are time-consuming and pose significant risks to personnel if problems arise. While infrared imaging can partially address the efficiency and safety issues, it requires additional infrared equipment, hindering large-scale deployment. In recent years, with the development of artificial intelligence and computer vision technologies, image-based methods for identifying defects in power equipment have become a research hotspot. For example, deep neural networks are used to identify cracks, corrosion, and damage in images, thereby assisting in maintenance decision-making.

[0197] Based on the above, the existing solutions have the following drawbacks in terms of fine-grained verification of substation equipment:

[0198] The model structure is complex, the inference cost is high, and it is difficult to deploy: For example, methods based on Faster R-CNN and improved YOLOv8s usually require large model size and computing resources, making it difficult to deploy and use in low computing power environments such as edge devices or industrial sites. They are not suitable for real-time online monitoring of equipment. At the same time, if an anomaly occurs and post-processing such as alarms is desired, for typical machine learning models, the output images require a supporting system for post-processing. However, the image and text multimodal model can immediately output the anomaly category after detecting an anomaly, without the need for a redundant post-processing system, which makes it easier to deploy.

[0199] The training process relies on a large number of manually labeled bounding boxes or labels, resulting in high data costs. Traditional machine learning methods, due to the availability of only single-modal information from images, typically rely on precisely labeled defect areas or clear category labels to achieve good recognition results. This leads to high costs for sample collection and labeling, and makes it difficult to flexibly expand in real-world scenarios with diverse equipment states and scarce samples. Furthermore, for larger equipment such as transformers and expanders in substations, the number of abnormal samples is even scarcer, placing higher demands on the experimental model. This experiment fine-tunes and optimizes the original model so that it can achieve good results using a small amount of data for training.

[0200] The existing methods are unable to combine text for multimodal understanding and weak interactive semantic understanding capabilities: most of them are trained based on object detection or classification models. The models can only output fixed defect labels, but they cannot combine natural language text for flexible and interpretable recognition. They lack a deep understanding of defect semantics, which limits the scalability and interpretability of the models. Moreover, there is currently no publicly available technology to combine text descriptions and images as input to the model for training and recognition. They lack the comprehensive utilization of multimodal information and cannot achieve an intuitive recognition method of "image input to text output".

[0201] The response speed and traceability are relatively weak: For traditional methods, machine learning models will identify anomalies and output detection results after they are detected, but the output is only a single image. The specific time of occurrence needs to be detected manually, which makes it difficult to locate quickly. At the same time, the speed of transmitting images from edge devices to the system is slow. In contrast, after identifying anomalies, multimodal models will immediately output the time of occurrence of the anomaly and the anomaly category corresponding to the current scene in text form, which has a faster transmission and response speed.

[0202] To address the above problems, this invention proposes a method for identifying defects in power equipment. Please refer to [link / reference]. Figure 3First, video stream information is acquired from the camera, and image information is input into the model frame by frame. Here, only image information is needed from the field. Then, the images are input into the improved image encoder of the model to obtain image features and map them into a vector space shared with the text. At the same time, this invention only detects anomalies such as capacitor bulging and expander overshoot, so the text content can be predefined manually. The pre-set text representing the state is input into the text enhancement and segmentation module, which expands the text according to the input text and the defined normal and abnormal category classification method to obtain more text representations. Then, all the text is input into the text encoder. Note that all operations for the text encoder input are predefined and do not need to be modified in the actual application stage. Then, the features output by the text encoder and the image encoder are compared to match the corresponding text for each input image. When the output text is found to be content representing predefined abnormal information, an alarm is generated, and the abnormality type and time information are output to the system to facilitate staff inspection and repair.

[0203] Compared with the prior art, the present invention has the following significant advantages: 1) It does not rely on target box annotation and complex post-processing, reduces data dependence and simplifies the usage process: Existing methods mostly rely on target detection box annotation or specific structural design to locate defect areas. The method of the present invention only needs the image and its corresponding natural language description as supervision signals to train the model, which greatly reduces the data construction cost and annotation work. Moreover, there is no need for post-processing steps such as detection box regression and non-maximum suppression in the inference stage, making the structure more concise and efficient.

[0204] 2) Applicable to scenarios with few samples: To address the problem of difficulty in collecting fault samples of power equipment, this invention introduces the LoRA (low-rank adaptation) module into the image encoder and text encoder of MobileCLIP, and adds a Prompt Tuning mechanism to the text encoder. This significantly improves the model's performance in scenarios with few samples while keeping parameter overhead extremely low. This is not yet reflected in existing technologies and is particularly suitable for the problem of scarce defect samples in industrial scenarios.

[0205] 3) Suitable for real-time operation of edge devices: Compared with traditional deep models, this invention uses MobileViT and MobileBERT as encoders, and the overall MobileCLIP model has fewer parameters and lower computational load. It has the ability to run in resource-constrained environments such as embedded devices and substation front-end hosts, meeting the actual needs of real-time detection and alarm on site.

[0206] 4) High scalability: This invention expands categories by introducing a templated text description method. There is no need to reconstruct the model structure or output dimensions, nor is a large amount of labeled data required. Only by supplementing the image-text pair samples or expanding the text set, new fault types can be quickly adapted, which greatly improves the long-term availability and maintenance efficiency of the system.

[0207] In summary, this invention differs from traditional object detection-based anomaly detection methods by introducing an image-text comparison model for industrial equipment anomaly identification. By constructing a pairing relationship between images and semantic descriptions, a lightweight image-text matching model is trained. During deployment, only an image needs to be input to output the most matching text description, and anomalies are determined based on this description. Furthermore, addressing the scarcity of anomaly sample data and the difficulty of small-sample training in real-world scenarios, this invention structurally enhances the image encoder and text encoder structures of MobileCLIP. LoRA modules are inserted into the Transformer modules of these two encoders to adapt to scene distribution with minimal parameter changes, enhancing the effectiveness of small-sample training. Finally, a trainable Prompt Embedding sequence is introduced before the text encoder to guide the text expression to be sensitive to specific fault keywords. These fine-tuning methods for the MobileCLIP model structure are another key aspect.

[0208] In this embodiment of the invention, a method for identifying defects in substation equipment is provided. This method acquires images of substation equipment and multiple predefined text descriptions of substation equipment status. A pre-set lightweight image-text matching model is used to identify defects in the substation equipment based on the images and the pre-defined text descriptions, and the identification result is output. Based on this scheme, by acquiring images of substation equipment and multiple predefined text descriptions of equipment status, and using a pre-set lightweight image-text matching model to perform correlation analysis between the two, the process of identifying and outputting results for substation equipment defects is achieved. This invention enables the identification process to utilize both the visual features of the images and the semantic information of the text, effectively improving the identification effect in real-world scenarios.

[0209] For better illustration, refer to Figure 4 The diagram illustrates the steps of training a pre-built lightweight image-text matching model according to Embodiment 2 of the present invention. This process may include the following steps:

[0210] Step 401: Obtain images of substation equipment for model training.

[0211] It should be noted that the substation equipment images used for model training are obtained from video stream data returned by substation monitoring cameras or from the internal substation database. However, the primary source of raw data is real-world video from substation monitoring cameras. For multiple substation scenarios, frames are periodically sampled from videos at different time periods to obtain static image frames. The frame extraction frequency can be flexibly adjusted based on the time granularity and scene stability. All acquired frames undergo manual or semi-automatic screening to ensure they contain clearly identifiable equipment areas. At this stage, corresponding task-specific labels are constructed. Then, image-text comparison data is built using methods such as manual recognition and cropping, and image description generation.

[0212] Step 402: Preprocess the substation equipment images and predefined substation equipment status text descriptions used for model training, and output image-text pairing.

[0213] It should be noted that the specific preprocessing process for the substation equipment images and predefined substation equipment status text descriptions used for model training is as follows: 1) Image screening and region cropping: The target region (capacitor or expander) in the acquired image is located and cropped, retaining only the equipment body region and removing irrelevant background information to improve the model's ability to perceive local abnormal features. Images are divided into two categories: normal samples and abnormal samples. Normal samples represent equipment images with intact structures and no abnormal signs, while abnormal samples indicate obvious bulging, overshooting, or other deformation phenomena in the equipment, which are manually judged as abnormal. 2) Text description annotation and enhancement: such as Figure 5 As shown, for each image, a core anomaly category label is first assigned manually or by a rule-based system, such as "capacitor bulging," "expander overshoot," or "equipment normal." Then, based on this core description, multiple text templates are designed for natural language reconstruction and expansion. This involves inputting the manually assigned label into the templates to automatically obtain additional descriptions. Multiple semantically equivalent but expressively diverse text descriptions are generated according to predefined categories, improving the model's robustness and generalization ability to language expression. For example, for "capacitor bulging," the following descriptions can be generated: "A capacitor bulging phenomenon was detected on the equipment's exterior," "The current equipment status is abnormal, and capacitor bulging is suspected," and "The capacitor bulging problem is quite obvious, and it is necessary to promptly investigate potential operational safety hazards." Each image will ultimately form multiple image-text pairs with its corresponding texts, i.e., image-text pairing, used for image-text matching or comparative learning during the training phase.

[0214] Step 403: Using the initial lightweight image-text matching model, the training similarity of the image-text pairing is output.

[0215] Step 404: Substitute the training similarity into the preset loss function and take the derivative to output the model gradient.

[0216] Step 405: Update the model parameters of the initial lightweight image-text matching model using model gradient, determine the intermediate lightweight image-text matching model, and count the number of model updates in real time.

[0217] Step 406: Determine whether the number of model updates has reached the preset number of training iterations.

[0218] Step 407: If so, use the intermediate lightweight image-text matching model as the trained pre-set lightweight image-text matching model.

[0219] The initial lightweight image-text matching model is the lightweight image-text matching model to be trained.

[0220] It should be noted that the preset loss function is the InfoNCE loss (Information Noise-Contrastive Estimation), and its expression is as follows:

[0221] ;

[0222] Where L is the loss value corresponding to the preset loss function; Here, N represents the temperature parameter; N is the number of negative samples. The feature corresponding to the i-th negative sample (the text feature of the negative sample).

[0223] Furthermore, if the number of model updates does not reach the preset number of training iterations, the intermediate lightweight image-text matching model is used as the new initial lightweight image-text matching model, and the process jumps to step 403 until the number of model updates reaches the preset number of training iterations. The intermediate lightweight image-text matching model determined when the number of model updates reaches the preset number of training iterations is then used as the trained preset lightweight image-text matching model.

[0224] In this embodiment of the invention, the model is trained by pairing images and text, allowing it to deeply integrate visual features and semantic information during the learning process, thereby enhancing its ability to understand the status of substation equipment. By leveraging loss functions and gradient update mechanisms, the model can accurately optimize parameters, improving its ability to distinguish between positive and negative samples. This effectively addresses the problems of weak generalization ability and poor adaptability to complex scenes caused by traditional models relying solely on image features, fundamentally solving the problem of unsatisfactory recognition results in real-world scenarios. The trained model thus possesses higher recognition accuracy and stability in practical applications.

[0225] Please see Figure 6 , Figure 6 This is a structural block diagram of a power equipment defect identification device provided in Embodiment 3 of the present invention.

[0226] The present invention provides a defect identification device for power equipment, comprising:

[0227] The acquisition module 601 is used to acquire images of substation equipment and multiple predefined text descriptions of substation equipment status.

[0228] The identification module 602 is used to identify substation equipment defects based on substation equipment images and multiple predefined substation equipment status text descriptions using a pre-set lightweight image-text matching model, and output the substation equipment defect identification results.

[0229] Furthermore, the pre-built lightweight image-text matching model includes a lightweight image encoder and a lightweight text encoder; the recognition module 602 includes:

[0230] The first submodule is used to encode images of substation equipment using a lightweight image encoder and output the image features corresponding to the substation equipment images.

[0231] The second submodule is used to encode the text descriptions of each predefined substation equipment status using a lightweight text encoder, and output the text features of each predefined substation equipment status text description.

[0232] The third submodule is used to identify substation equipment defects based on image features and various text features, and output the substation equipment defect identification results.

[0233] Furthermore, the lightweight image encoder includes a main intervention processing layer, an inverse residual module, a low-rank adaptive image module with multiple attention mechanisms, and an image projection head; the first submodule includes:

[0234] The first unit is used to preprocess images of substation equipment and output images that conform to the input dimensions of the model.

[0235] The second unit is used to perform preliminary downsampling on images that conform to the model input size through the main intervention processing layer, and generate downsampled image features.

[0236] The third unit is used to take downsampled image features as input to the inverse residual module and output inverse residual image features;

[0237] The fourth unit is used to perform convolution operations on the inverse residual image features and output local image features;

[0238] The fifth unit is used to perform iterative image enhancement on local image features using multiple attention mechanism low-rank adaptive image modules to generate global image features;

[0239] The sixth unit is used to stitch together global and local image features, output stitched image features, and perform convolution operations on the stitched image features to output target fused image features;

[0240] The seventh unit is used to perform global pooling on the target fused image features, output the target global pooled image features, and extract vectors from the target global pooled image features to generate image representation vectors;

[0241] The eighth unit is used to project the image representation vector onto the image using an image projection head to generate image features corresponding to the substation equipment images.

[0242] Furthermore, the second unit is specifically used for:

[0243] Channel expansion is performed on the downsampled image features to generate expanded image features;

[0244] Perform depthwise convolution on the extended image features to output the depthwise convolution image features;

[0245] Perform a projection operation on the features of the deep convolutional image to generate projected image features;

[0246] The features of the projected image and the downsampled image are added together to output the inverse residual image features.

[0247] Optionally, the low-rank adaptive image module of the attention mechanism includes a multi-head attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning; the feature processing procedure of the low-rank adaptive image module of the attention mechanism is as follows:

[0248] The input image features to the low-rank adaptive image module of the attention mechanism are divided into blocks, and multiple image sub-features are output.

[0249] Each image sub-feature is flattened to generate multiple image sub-feature vectors;

[0250] Add positional encoding to each image sub-feature vector to output multiple image sub-feature codes;

[0251] A multi-head attention module based on low-rank matrix fine-tuning is used to output the multi-head attention image features corresponding to each image sub-feature encoding according to the sub-feature encoding of each image.

[0252] Each multi-head attention image feature is used as input to a feedforward network based on low-rank matrix fine-tuning, and the output is the feedforward image feature corresponding to each multi-head attention image feature.

[0253] The features of each feedforward image are averaged and aggregated to output the aggregated image features.

[0254] Embedding projection is performed on the aggregated image features to generate output image features.

[0255] Furthermore, the lightweight text encoder includes multiple attention mechanisms, a low-rank adaptive text module, and a tagging convergence layer; the second submodule is specifically used for:

[0256] Each predefined substation equipment status text description is preprocessed and output as a word sequence corresponding to the predefined substation equipment status text description.

[0257] Each word sequence is concatenated with a pre-set learnable prompt to output the target word sequence corresponding to each word sequence;

[0258] A low-rank adaptive text module employing multiple attention mechanisms performs iterative semantic enhancement based on each target word sequence, outputting word sequence features corresponding to each target word sequence;

[0259] Each word sequence feature is subjected to layer normalization to generate word sequence normalization features corresponding to each word sequence feature;

[0260] The normalized features of each word sequence are input into the label aggregation layer, and the text features corresponding to the normalized features of each word sequence are output.

[0261] Optionally, the attention mechanism low-rank adaptive text module includes a self-attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning; the feature processing procedure of the attention mechanism low-rank adaptive text module is as follows:

[0262] The input text features of the low-rank adaptive text module of the input attention mechanism are processed by the self-attention module based on low-rank matrix fine-tuning, and the self-attention text features corresponding to the input text features are output.

[0263] A feedforward network based on low-rank matrix fine-tuning is used to perform nonlinear transformation on self-attention text features and output feedforward text features.

[0264] The feedforward text features and the input text features are added together to output the added text features. Then, the added text features are layer normalized to generate the output text features.

[0265] Furthermore, the third submodule is specifically used for:

[0266] The similarity between the image features and each text feature is calculated, and the similarity between the image features and each text feature is output.

[0267] The predefined substation equipment status text description associated with the text feature corresponding to the highest similarity is selected as the substation equipment defect identification result.

[0268] In one optional device embodiment, it further includes:

[0269] The first module is used to acquire images of substation equipment for model training;

[0270] The second module is used to preprocess the images of substation equipment used for model training and the predefined text descriptions of substation equipment status, and output image-text pairing.

[0271] The third module is used to use the initial lightweight image-text matching model to pair images and text, and output the training similarity corresponding to the image-text pairing.

[0272] The fourth module is used to substitute the training similarity into the preset loss function and calculate the derivative to output the model gradient;

[0273] The fifth module is used to update the model parameters of the initial lightweight image-text matching model using model gradients, determine the intermediate lightweight image-text matching model, and count the number of model updates in real time.

[0274] The sixth module is used to determine whether the number of model updates has reached the preset number of training iterations;

[0275] The seventh module is used to use the intermediate lightweight image-text matching model as the pre-trained lightweight image-text matching model if the condition is met.

[0276] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, modules, sub-modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0277] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0278] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0279] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying defects in power equipment, characterized in that, include: Acquire images of substation equipment and multiple predefined text descriptions of substation equipment status; A pre-built lightweight image-text matching model is used to identify substation equipment defects based on the substation equipment images and multiple predefined substation equipment status text descriptions, and the substation equipment defect identification results are output.

2. The method for identifying defects in power equipment according to claim 1, characterized in that, The pre-set lightweight image-text matching model includes a lightweight image encoder and a lightweight text encoder; the pre-set lightweight image-text matching model is used to identify substation equipment defects based on the substation equipment images and multiple predefined substation equipment status text descriptions, and outputs substation equipment defect identification results, including: The lightweight image encoder is used to encode the substation equipment image and output the image features corresponding to the substation equipment image. The lightweight text encoder is used to encode the text descriptions of the predefined substation equipment status respectively, and outputs the text features of the text descriptions of the predefined substation equipment status. Based on the image features and the text features, substation equipment defects are identified, and the substation equipment defect identification results are output.

3. The method for identifying defects in power equipment according to claim 2, characterized in that, The lightweight image encoder includes a main intervention processing layer, an inverse residual module, multiple attention mechanism low-rank adaptive image modules, and an image projection head; The step of using a lightweight image encoder to encode the substation equipment image and outputting the image features corresponding to the substation equipment image includes: The images of the substation equipment are preprocessed to output images that conform to the input dimensions of the model; The main intervention processing layer performs preliminary downsampling on the image that conforms to the model input size to generate downsampled image features. The downsampled image features are used as input to the inverse residual module, and the inverse residual image features are output. Perform convolution operation on the inverted residual image features to output local image features; Multiple attention-based low-rank adaptive image modules are used to iteratively enhance the local image features to generate global image features; The global image features and the local image features are concatenated to output the concatenated image features, and the concatenated image features are convolved to output the target fused image features. The target fused image features are globally pooled to output the target globally pooled image features, and vector extraction is performed on the target globally pooled image features to generate an image representation vector; The image representation vector is projected using an image projection head to generate image features corresponding to the substation equipment image.

4. The method for identifying defects in power equipment according to claim 3, characterized in that, The step of using the downsampled image features as input to the inverse residual module and outputting inverse residual image features includes: The downsampled image features are channel-expanded to generate expanded image features; Perform depthwise convolution on the extended image features to output depthwise convolutional image features; The deep convolutional image features are projected to generate projected image features; The projected image features and the downsampled image features are added together to output the inverse residual image features.

5. The method for identifying defects in power equipment according to claim 3, characterized in that, The attention mechanism low-rank adaptive image module includes a multi-head attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning; the feature processing procedure of the attention mechanism low-rank adaptive image module is as follows: The input image features input to the low-rank adaptive image module of the attention mechanism are divided into blocks, and multiple image sub-features are output. Each of the image sub-features is flattened to generate multiple image sub-feature vectors; Add positional encoding to each of the image sub-feature vectors to output multiple image sub-feature codes; A multi-head attention module based on low-rank matrix fine-tuning is used to output the multi-head attention image features corresponding to each image sub-feature encoding according to the image sub-feature encoding; Each of the multi-head attention image features is used as input to a feedforward network based on low-rank matrix fine-tuning, and the corresponding feedforward image features are output. The feedforward image features are averaged and aggregated to output the aggregated image features. The aggregated image features are embedded and projected to generate output image features.

6. The method for identifying defects in power equipment according to claim 2, characterized in that, The lightweight text encoder includes multiple attention-based low-rank adaptive text modules and a tagging and convergence layer. The lightweight text encoder encodes the text of each predefined substation equipment status description, outputting the text features of each predefined substation equipment status description, including: Each of the predefined substation equipment status text descriptions is preprocessed to output the word sequence corresponding to each of the predefined substation equipment status text descriptions; Each of the aforementioned word sequences is concatenated with a preset learnable prompt to output the target word sequence corresponding to each of the aforementioned word sequences; A low-rank adaptive text module employing multiple attention mechanisms performs iterative semantic enhancement based on each target word sequence, outputting word sequence features corresponding to each target word sequence; Each of the word sequence features is subjected to layer normalization to generate word sequence normalization features corresponding to each of the word sequence features; The normalized features of each word sequence are input into the labeling convergence layer, and the text features corresponding to each normalized feature of the word sequence are output.

7. The method for identifying defects in power equipment according to claim 6, characterized in that, The attention mechanism low-rank adaptive text module includes a self-attention module based on low-rank matrix fine-tuning and a feedforward network based on low-rank matrix fine-tuning; the feature processing procedure of the attention mechanism low-rank adaptive text module is as follows: The input text features input to the low-rank adaptive text module of the attention mechanism are processed by a self-attention module based on low-rank matrix fine-tuning to calculate attention and output the self-attention text features corresponding to the input text features. A feedforward network based on low-rank matrix fine-tuning is used to perform a nonlinear transformation on the self-attention text features, and the feedforward text features are output. The feedforward text features and the input text features are added together to output the text addition features. The text addition features are then layer-normalized to generate the output text features.

8. The method for identifying defects in power equipment according to claim 2, characterized in that, The process of identifying substation equipment defects based on the image features and each of the text features, and outputting the substation equipment defect identification result, includes: The similarity between the image features and each of the text features is calculated, and the similarity between the image features and each of the text features is output. The predefined substation equipment status text description associated with the text feature corresponding to the highest similarity is selected as the substation equipment defect identification result.

9. The method for identifying defects in power equipment according to claim 1, characterized in that, The training process of the pre-built lightweight image-text matching model is as follows: Acquire images of substation equipment for model training; The images of substation equipment used for model training and the predefined text descriptions of substation equipment status are preprocessed to output image-text pairings. An initial lightweight image-text matching model is used to pair the image and text, and the training similarity corresponding to the image-text pair is output. Substitute the training similarity into the preset loss function and take the derivative to output the model gradient; The model parameters of the initial lightweight image-text matching model are updated using the model gradient to determine the intermediate lightweight image-text matching model, and the number of model updates is counted in real time. Determine whether the number of model updates has reached the preset number of training iterations; If so, the intermediate lightweight image-text matching model is used as the trained preset lightweight image-text matching model.

10. A defect identification device for power equipment, characterized in that, include: The acquisition module is used to acquire images of substation equipment and multiple predefined text descriptions of substation equipment status; The identification module is used to identify substation equipment defects based on the substation equipment images and multiple predefined substation equipment status text descriptions using a pre-set lightweight image-text matching model, and output the substation equipment defect identification results.

Citation Information

Cited By

  • Power generation equipment defect identification method and system based on open set and multi-modal model

    CN122049606A