Substation equipment identification method, device and equipment and readable storage medium
By combining the visual network and language encoding convolutional layer in the improved CLIP model for substation equipment identification, the problems of low efficiency and high false detection rate in traditional methods are solved, and efficient and accurate equipment identification and timely detection of abnormal situations are achieved.
Patent Information
- Application Number
- CN202510784090.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional substation equipment type verification methods are inefficient and have a high false positive rate. They mainly rely on manual inspections and are easily affected by human experience.
An improved CLIP model is adopted. By adding an Adapter module to the visual network and replacing the language network with a language encoding convolutional layer, the language encoding convolutional layer is used to inject the device language representation, and the visual network and the language encoding convolutional layer are combined to perform substation equipment recognition.
It improves the accuracy and efficiency of substation equipment identification, reduces the false detection rate, and can promptly detect abnormal conditions in image acquisition to ensure safe and stable power supply.
Smart Images

Figure CN120635668A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more specifically, to a method, apparatus, device, and readable storage medium for identifying substation equipment. Background Art
[0002] In power systems, accurate verification of substation equipment types is crucial for ensuring safe and stable operation. With the continuous expansion and increasing intelligence of power systems, higher requirements are placed on the accuracy and efficiency of substation equipment type verification. However, traditional methods for verifying substation equipment types rely primarily on manual inspections, which is not only inefficient but also susceptible to human experience, resulting in a high rate of false positives. Summary of the Invention
[0003] In view of this, the present application provides a substation equipment identification method, apparatus, device and readable storage medium to address the shortcomings of low efficiency and high false detection rate in existing substation equipment type verification technology.
[0004] In order to achieve the above objectives, the following solutions are proposed:
[0005] A method for identifying substation equipment, comprising:
[0006] Obtaining a substation image and an improved CLIP model, wherein an adapter module is added to each Vision Transformer module of the vision network and the language network is replaced with a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation devices;
[0007] The improved CLIP model is used to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolutional layer is used to process the image features to determine the equipment identification of the substation equipment area; if not, a reminder is generated to indicate that the substation image does not contain substation equipment.
[0008] Optionally, obtain an improved CLIP model, including:
[0009] Build a CLIP model that includes a visual network and a language network;
[0010] In the CLIP model, an Adapter module is added after the multi-head self-attention layer of each Vision Transformer module in the vision network. At least two MLP modules are added before the input layer of the language network in the CLIP model. Each MLP module consists of a three-layer linear layer mixed with a Layer Norm module with a residual branch to form the initial CLIP model.
[0011] Acquire multiple training samples, each training sample comprising a training device image and a corresponding training device code, wherein the training device images have the same size;
[0012] Using each training sample, the initial CLIP model is trained until the initial CLIP model meets a preset stopping condition;
[0013] The language network of the final initial CLIP model is updated to a language encoding convolutional layer to generate an improved CLIP model.
[0014] Optionally, obtaining multiple training samples includes:
[0015] Obtain different training images marked with training device image areas and their corresponding training device identifiers, where all training device identifiers of each training image constitute identifiers corresponding to all substation devices;
[0016] Extracting a training device image region from each training image and normalizing the size of each training device image region to obtain each training device image;
[0017] Using the same coding template, based on each training device identifier, obtain the training device code corresponding to each training device identifier;
[0018] Each training device code is used as the annotation label of the corresponding training device image to form a training sample.
[0019] Optionally, updating the language network of the final initial CLIP model to a language encoding convolutional layer to generate an improved CLIP model includes:
[0020] In combination with the coding template, based on the identification of each substation equipment, a language description corresponding to each substation equipment is generated;
[0021] Perform dimension processing and numerical scaling on each language description to generate a device language representation for each substation device;
[0022] Initialize the language network as a convolutional layer with a kernel size of 1x1;
[0023] The conv_weight unit in the convolutional layer is replaced with each device language representation, and the conv_bias unit is replaced with a logit_bias unit to form a language encoding convolutional layer, thereby obtaining an improved CLIP model.
[0024] Optionally, the training of the initial CLIP model using various training samples includes:
[0025] In each iterative training, each training sample is sequentially input into the initial CLIP model to obtain the prediction result of the initial CLIP model;
[0026] Calculate the prediction loss value based on each training sample and its corresponding prediction result;
[0027] Based on the predicted loss value, parameters of each Adapter module and each MLP module are adjusted.
[0028] Optionally, the calculation of the prediction loss value based on each training sample and its corresponding prediction result includes:
[0029] Obtain the image training features and text training features extracted by the initial CLIP model based on each training sample and calculate the similarity matrix;
[0030] In combination with the Sigmoid loss function, the prediction loss value is calculated based on each training sample and its corresponding prediction result, and the similarity matrix.
[0031] Optionally, after processing the image features using the language encoding convolution layer to determine the equipment identification of the substation equipment area, the method further includes:
[0032] Based on the improved CLIP model, a device language representation corresponding to the device identifier is obtained, and a similarity between the corresponding device language representation and the image feature is calculated. When the similarity exceeds a similarity threshold, the device identifier is output; when the similarity is less than the similarity threshold, a step of generating a reminder indicating that the substation image does not contain substation equipment is executed.
[0033] A substation equipment identification device, comprising:
[0034] An acquisition module, configured to acquire substation images and an improved CLIP model, wherein an adapter module is added to each Vision Transformer module of the visual network and the language network is replaced with a language encoding convolutional layer, which is injected with the device language representation corresponding to all substation devices;
[0035] an identification module for using the improved CLIP model to identify whether a substation equipment area exists in the substation image; if so, generating image features corresponding to the substation equipment area using the visual network, and processing the image features using the language encoding convolutional layer to determine the equipment identification of the substation equipment area; if not, generating a reminder indicating that the substation image does not contain substation equipment.
[0036] A substation equipment identification device, comprising a memory and a processor;
[0037] The memory is used to store programs;
[0038] The processor is used to execute the program to implement each step of the above-mentioned substation equipment identification method.
[0039] A readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above-mentioned substation equipment identification method are implemented.
[0040] It can be seen from the above technical solutions that the substation equipment recognition method provided by the present application can obtain substation images and an improved CLIP model. In the improved CLIP model, an Adapter module is added to each VisionTransformer module of the visual network and the language network is replaced by a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation equipment; based on this, the improved CLIP model of the present application can retain the general feature extraction capability of the original visual network while using multiple Adapter modules to adjust the various features extracted by the visual network to match the visual network with the substation equipment recognition scene; while ensuring the feature extraction capability of the visual network, the matching degree between the visual network and each substation is improved; at the same time, the language encoding convolutional layer is simpler than the original language network structure, which reduces the amount of calculation in the recognition process and improves the recognition efficiency. At the same time, due to the language encoding convolutional layer The device language representation of all substation equipment is included, and the features collected by the visual network can be matched with the language representation of each device to reduce the false detection rate. On this basis, the present application can use the improved CLIP model to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolution layer is used to process the image features to determine the device identification of the substation equipment area. If not, a reminder is generated to indicate that the substation image does not contain substation equipment. Based on this, the present application can use the visual network and language encoding convolution layer in the improved CLIP model for integrated recognition, reduce the recognition delay caused by data transmission, and improve the recognition efficiency while reducing the false detection rate. At the same time, the present application can not only generate device identification but also generate reminders so that users can promptly discover abnormal conditions in image acquisition, further accelerating the substation equipment verification process. It can be seen that the present application can use the improved CLIP model to identify substation equipment, while improving recognition accuracy, accelerating recognition efficiency, and ensuring safe and stable power supply. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0042] Figure 1 This is a flow chart of a method for identifying substation equipment disclosed in an embodiment of the present application;
[0043] Figure 2 This is a structural block diagram of a substation equipment identification device disclosed in an embodiment of the present application;
[0044] Figure 3 This is a hardware structure block diagram of a substation equipment identification device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0046] An embodiment of the present application provides a substation equipment identification method, which can be applied to various image processing systems or substation management systems, and can also be applied to various computer terminals or smart terminals. Its execution subject can be the processor or server of the computer terminal or smart terminal.
[0047] Next, combine Figure 1 The substation equipment identification method of the present application is introduced in detail, including the following steps:
[0048] Step S1: Obtain a substation image and an improved CLIP model.
[0049] Specifically, the substation image captured by the smart camera configured in the substation can be obtained.
[0050] By adjusting the shooting angle of the smart camera, substation images in different areas of the same substation can be collected.
[0051] In the improved CLIP model, an Adapter module is added to each Vision Transformer module of the vision network, and the language network is replaced by a language encoding convolutional layer, which is injected with the device language representation corresponding to all substation equipment.
[0052] Among them, the CLIP model is a contrastive language-image pretraining model, which can perform well in cross-domain and cross-modal open-world representation and is the basis for various visual and multimodal tasks.
[0053] The Adapter module is an adapter module that can be inserted into an existing model. By performing operations such as dimensionality reduction, information exchange, and dimensionality increase on the input features, it enables the model to adapt to specific tasks without changing the main structure of the original model, thereby enhancing the model's ability to learn new tasks.
[0054] The Vision Transformer module is a visual processing component based on the Transformer architecture. Its core is to divide the image into patches of fixed size, convert it into a sequence, capture global feature dependencies through a multi-head self-attention mechanism, and then optimize local features through a feedforward neural network.
[0055] The language encoding convolutional layer belongs to the convolutional layer, which records the language representation of each device; it can be used to perform similarity comparison with the features collected by the visual network to evaluate whether the substation equipment identified by the visual network is correct.
[0056] Step S2: using the improved CLIP model to identify whether there is a substation equipment area in the substation image; if so, using the visual network to generate image features corresponding to the substation equipment area, and using the language encoding convolutional layer to process the image features to determine the equipment identification of the substation equipment area; if not, generating a reminder indicating that the substation image does not contain substation equipment.
[0057] Specifically, a substation image can be input into the improved CLIP model. A visual network is used to generate a target box marking the substation equipment area. Feature extraction is then performed on the area within the target box to generate image features. A language encoding convolutional layer is then used to match and identify these image features, determining the corresponding image features. Similarity is then determined between these features to determine whether the target box was generated accurately. If not, the improved CLIP model outputs a reminder indicating that the substation image does not contain substation equipment.
[0058] It can be seen from the above technical solutions that the substation equipment identification method provided by the present application can obtain substation images and an improved CLIP model. In the improved CLIP model, an Adapter module is added before each VisionTransformer module of the visual network and the language network is replaced by a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation equipment. Based on this, the improved CLIP model of the present application can retain the general feature extraction capability of the original visual network while using multiple Adapter modules to adjust the various features extracted by the visual network to match the visual network with the substation equipment identification scene; while ensuring the feature extraction capability of the visual network, the matching degree between the visual network and each substation is improved; at the same time, the language encoding convolutional layer is simpler than the original language network structure, which reduces the amount of calculation in the recognition process and improves the recognition efficiency. At the same time, due to the language encoding convolutional layer The device language representation of all substation equipment is included, and the features collected by the visual network can be matched with the language representation of each device to reduce the false detection rate. On this basis, the present application can use the improved CLIP model to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolution layer is used to process the image features to determine the device identification of the substation equipment area. If not, a reminder is generated to indicate that the substation image does not contain substation equipment. Based on this, the present application can use the visual network and language encoding convolution layer in the improved CLIP model for integrated recognition, reduce the recognition delay caused by data transmission, and improve the recognition efficiency while reducing the false detection rate. At the same time, the present application can not only generate device identification but also generate reminders so that users can promptly discover abnormal conditions in image acquisition, further accelerating the substation equipment verification process. It can be seen that the present application can use the improved CLIP model to identify substation equipment, while improving recognition accuracy, accelerating recognition efficiency, and ensuring safe and stable power supply.
[0059] In some embodiments of the present application, the process of obtaining the improved CLIP model in step S1 is described in detail, and the steps are as follows:
[0060] S10. Construct a CLIP model that includes a visual network and a language network.
[0061] Specifically, a CLIP model may be generated, wherein the CLIP model includes a visual network and a language network.
[0062] Among them, the visual network can include 4 Vision Transformer modules.
[0063] The Vision Transformer module can contain multi-head self-attention layers, residual connections and layer normalization, and feed-forward neural network layers.
[0064] S11. Add an Adapter module after the multi-head self-attention layer of each Vision Transformer module in the vision network of the CLIP model, and add at least two MLP modules before the input layer of the language network in the CLIP model. Each MLP module consists of a three-layer linear layer mixed Layer Norm module with a residual branch to form the initial CLIP model.
[0065] Specifically, an Adapter module can be added after the multi-head self-attention layer in each Vision Transformer module in the CLIP model.
[0066] Among them, the Linear layer of the Adapter module can be used for dimensionality reduction, the deep separation convolutional layer Conv can be used to enhance the encoding after dimensionality reduction and realize information interaction. GELU and Linear can process the dimensionality-reduced encoding, regress the initial feature dimension, and obtain and output the encoding features to the main branch of the visual network.
[0067] The Layer Normalization module is a normalization technique in deep learning that accelerates training and improves stability by normalizing the features of a single sample.
[0068] The MLP module is a multi-layer perceptron module (MLP module for short). It can implement complex function mapping by performing multiple nonlinear transformations on the input data.
[0069] S12. Acquire multiple training samples, each training sample comprising a training device image and a corresponding training device code, wherein the sizes of the training device images are the same.
[0070] Specifically, the training device codes corresponding to all training samples may correspond to all substation devices.
[0071] Different training samples may correspond to the same substation equipment and have the same training equipment code.
[0072] The size of each training device image can be set according to the input layer of the CLIP model.
[0073] S13. Utilize various training samples to train the initial CLIP model until the initial CLIP model meets a preset stopping condition.
[0074] Specifically, the initial CLIP model can be iteratively trained multiple times using each training sample, and the parameters of the four Adapter modules and each MLP module of the initial CLIP model can be adjusted until the initial CLIP model converges or the number of iterations reaches a preset iteration threshold.
[0075] S14. Update the language network of the final initial CLIP model to a language encoding convolutional layer to generate an improved CLIP model.
[0076] Specifically, the language network of the trained initial CLIP model can be compressed into a language encoding convolutional layer, and the improved CLIP model is obtained after compression.
[0077] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining an improved CLIP model. Through the above method, the CLIP model after module adjustment can be trained to improve the feature extraction capability of the visual network while maintaining the feature extraction capability of the visual network. While maintaining the text feature extraction capability of the language network, the computational cost of the finally generated improved CLIP model can be reduced.
[0078] In some embodiments of the present application, the process of obtaining multiple training samples in step S12, each training sample including a training device image and a corresponding training device code, wherein the training device images have the same size, is described in detail as follows:
[0079] S120: Acquire different training images marked with training device image areas and their corresponding training device identifiers, where all training device identifiers of the various training images constitute identifiers corresponding to all substation equipment.
[0080] Specifically, smart cameras or drones can be used to collect training images from different training stations.
[0081] You can also use open source databases to obtain multiple training images.
[0082] The same training image can contain multiple training device image regions marked with bounding boxes.
[0083] Different training device image regions in the same training image may correspond to the same training device identifier.
[0084] S121 . Extracting a training device image region from each training image, and normalizing the size of each training device image region to obtain each training device image.
[0085] Specifically, each training device image region may be cropped from each training image;
[0086] When the training device image area is smaller than the training size, a white canvas consistent with the training size can be generated, and the training device image area is pasted to the center area of the white canvas to obtain the training device image.
[0087] When the image area of the training device is larger than the training size, the long side of the image area of the training device is scaled and the short side pixels are filled to obtain the training device image.
[0088] S122. Using the same coding template, based on each training device identifier, obtain the training device code corresponding to each training device identifier.
[0089] Specifically, the coding template may include multiple descriptions of the image, and the training device identifier is written into the coding template as a filler to obtain the training device code.
[0090] For example, the encoding template could look like this:
[0091] "a photo of a {object}"
[0092] "a picture of a {object}"
[0093] "an image of a {object}"
[0094] "a close-up of a {object}"
[0095] "a high-resolution photo of a {object}"
[0096] "a image of a {object} with white background"
[0097] You can replace {object} with the training device identifier to obtain the training device code.
[0098] S123: Use each training device code as a label for the corresponding training device image to form a training sample.
[0099] Specifically, the training device code and the training device image may be aligned in sequence to obtain a training sample.
[0100] It can be seen from the above technical solution that this embodiment provides an optional method for obtaining multiple training samples. Through the above method, the size of each training sample and the text label expression method can be unified, thereby reducing the loss value caused by different expression methods during training and accelerating the training process.
[0101] In some embodiments of the present application, step S14, in which the language network of the final initial CLIP model is updated to a language encoding convolutional layer to generate an improved CLIP model, is described in detail. The steps are as follows:
[0102] S140 . Generate a language description corresponding to each substation device based on the identification of each substation device in combination with the coding template.
[0103] Specifically, the identification of each substation device may be written into a coding template to obtain a language description corresponding to each substation device.
[0104] S141. Perform dimension processing and numerical scaling on each language description to generate a device language representation of each substation device.
[0105] Specifically, each language description may be reshaped to convert the two-dimensional language description into a four-dimensional language description.
[0106] The logit_scale unit is used to numerically scale the four-dimensional language description to obtain the device language representation of the substation equipment.
[0107] S142. Initialize the language network as a convolutional layer with a kernel size of 1x1.
[0108] Specifically, the language network can be initialized to obtain a convolutional layer with a kernel size of 1x1.
[0109] S143. Replace the conv_weight unit in the convolutional layer with the language representation of each device, and replace the conv_bias unit with the logit_bias unit to form a language encoding convolutional layer, thereby obtaining an improved CLIP model.
[0110] Specifically, the conv_weight unit in the convolutional layer can be replaced with the language representation of each device, and the conv_bias unit can be replaced with the logit_bias unit to obtain the language encoding convolutional layer.
[0111] It can be seen from the above technical solution that this embodiment provides an optional method of replacing the language network with a language encoding convolutional layer. Through the above method, the language network can be better compressed while improving the matching degree between the language encoding convolutional layer and the substation equipment recognition.
[0112] In some embodiments of the present application, the process of step S13, using various training samples to train the initial CLIP model until the initial CLIP model meets a preset stopping condition, is described in detail. The steps are as follows:
[0113] S130 . During each iterative training, each training sample is sequentially input into the initial CLIP model to obtain a prediction result of the initial CLIP model.
[0114] Specifically, during each iterative training, the initial CLIP model can be used to analyze each training sample to obtain a prediction result corresponding to each training sample.
[0115] S131. Calculate the prediction loss value based on each training sample and its corresponding prediction result.
[0116] Specifically, the prediction loss value corresponding to each iterative training process can be calculated based on each training sample and its corresponding prediction result, the image training features and text training features generated by the initial CLIP model.
[0117] S132. Based on the predicted loss value, adjust the parameters of each Adapter module and each MLP module.
[0118] Specifically, you can refer to the predicted loss value and combine the gradient descent method to adjust the parameters of the four Adapter modules and each MLP module.
[0119] As can be seen from the above technical solution, this embodiment provides an optional method for iteratively training the initial CLIP model. Through the above method, only the parameters of each Adapter module and each MLP module can be adjusted to maintain the feature extraction and language analysis capabilities of the CLIP model.
[0120] In some embodiments of the present application, step S131, the process of calculating the prediction loss value based on each training sample and its corresponding prediction result, is described in detail. The steps are as follows:
[0121] S1310 , obtaining the image training features and text training features extracted from each training sample by the initial CLIP model, and calculating a similarity matrix.
[0122] Specifically, the similarity between any image training feature and any text training feature can be calculated, and each similarity constitutes a similarity matrix.
[0123] S1311. In combination with the Sigmoid loss function, based on each training sample and its corresponding prediction result, and the similarity matrix, calculate the prediction loss value.
[0124] Specifically, the Sigmoid loss function can be used to set the annotation label of the similarity matrix based on each training sample and its corresponding prediction result, and calculate the prediction loss value.
[0125] It can be seen from the above technical solution that this embodiment provides an optional method for calculating the prediction loss value. Through the above method, the Sigmoid loss function can be used to reduce computing consumption and accelerate the training process.
[0126] In some embodiments of the present application, considering the possibility of false detection in the visual network, a detection process can be added after the language encoding convolution layer is used to process the image features in step S2 to determine the equipment identification of the substation equipment area to improve the reliability of substation verification. The detection process will be described in detail below. The steps are as follows:
[0127] S20. Based on the improved CLIP model, obtain a device language representation corresponding to the device identifier, calculate the similarity between the corresponding device language representation and the image feature, and when the similarity exceeds a similarity threshold, output the device identifier; when the similarity is less than the similarity threshold, enter the step of generating a reminder indicating that the substation image does not contain substation equipment.
[0128] Specifically, each device language representation may correspond one-to-one to an identifier corresponding to each substation device.
[0129] The device language representation that matches the device identifier can be retrieved, and the similarity between the matching device language representation and the image features can be calculated;
[0130] It can be determined whether the similarity is greater than the similarity threshold;
[0131] If so, it can be determined that the target frame is generated correctly and the device identification is output;
[0132] If not, it can be determined that the target frame is generated incorrectly and a reminder is generated.
[0133] It can be seen from the above technical solution that compared with the previous embodiment, this embodiment adds a new detection process. Through the above process, the language representation of each device in the language encoding convolution layer can be used to verify the visual network, thereby improving the recognition reliability of this application.
[0134] Next, we will combine Figure 2 The substation equipment identification device provided in this application is introduced in detail. The substation equipment identification device provided below can be compared with the substation equipment identification method provided above.
[0135] See also Figure 2 It can be found that the substation equipment identification device may include:
[0136] Acquisition module 10, for acquiring substation images and an improved comparative language-image pre-trained CLIP model. In the improved CLIP model, an adapter module is added to each Vision Transformer module of the vision network, and the language network is replaced with a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation devices;
[0137] The recognition module 20 is used to use the improved CLIP model to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolution layer is used to process the image features to determine the equipment identification of the substation equipment area; if not, a reminder is generated to indicate that the substation image does not contain substation equipment.
[0138] Furthermore, the acquisition module 10 may include:
[0139] Visual network construction unit, used to build the CLIP model including visual network and language network;
[0140] The visual network transformation unit is used to add an Adapter module after the multi-head self-attention layer of each Vision Transformer module in the visual network of the CLIP model, and to add at least two MLP modules before the input layer of the language network in the CLIP model. Each MLP module consists of a three-layer linear layer mixed with a Layer Norm module with a residual branch to form the initial CLIP model;
[0141] A training sample acquisition unit is used to acquire a plurality of training samples, each training sample comprising a training device image and a corresponding training device code, wherein the sizes of the training device images are the same;
[0142] A model training unit, configured to train the initial CLIP model using each training sample until the initial CLIP model meets a preset stopping condition;
[0143] The language network transformation unit is used to update the language network of the final initial CLIP model into a language encoding convolutional layer to generate an improved CLIP model.
[0144] Furthermore, the training sample acquisition unit may include:
[0145] The first training sample acquisition subunit is used to obtain different training images marked with training equipment image areas and their corresponding training equipment identifiers, and all training equipment identifiers of each training image constitute the identifiers corresponding to all substation equipment;
[0146] The second training sample acquisition subunit is used to extract the training device image region from each training image and normalize the size of each training device image region to obtain each training device image;
[0147] The third training sample acquisition subunit is configured to obtain a training device code corresponding to each training device identifier based on each training device identifier using the same encoding template;
[0148] The fourth training sample acquisition subunit is configured to encode each training device as a label for the corresponding training device image to form a training sample.
[0149] Furthermore, the language network transformation unit may include:
[0150] A first language network transformation subunit is configured to generate a language description corresponding to each substation device based on the identification of each substation device in combination with the coding template;
[0151] The second language network transformation subunit is used to perform dimension processing and numerical scaling on each language description to generate a device language representation for each substation device;
[0152] The third language network transformation subunit is used to initialize the language network into a convolutional layer with a kernel size of 1x1;
[0153] The fourth language network transformation subunit is used to replace the conv_weight unit in the convolutional layer with each device language representation, and replace the conv_bias unit with the logit_bias unit to form a language encoding convolutional layer, thereby obtaining an improved CLIP model.
[0154] Furthermore, the model training unit may include:
[0155] A prediction result generating subunit, configured to input each training sample into the initial CLIP model in sequence during each iterative training to obtain a prediction result of the initial CLIP model;
[0156] A prediction loss value calculation subunit, used to calculate the prediction loss value based on each training sample and its corresponding prediction result;
[0157] The parameter adjustment subunit is used to adjust the parameters of each Adapter module and each MLP module based on the predicted loss value.
[0158] Furthermore, the predicted loss value calculation subunit may include:
[0159] The first prediction loss value calculation component is used to obtain the image training features and text training features extracted by the initial CLIP model based on each training sample and calculate the similarity matrix;
[0160] The second prediction loss value calculation component is used to calculate the prediction loss value based on each training sample and its corresponding prediction result and the similarity matrix in combination with the Sigmoid loss function.
[0161] Furthermore, the identification module 20 may further include:
[0162] a similarity calculation unit, configured to obtain, based on the improved CLIP model, a device language representation corresponding to the device identifier, calculate a similarity between the corresponding device language representation and the image feature, and output the device identifier when the similarity exceeds a similarity threshold; and generate a reminder indicating that the substation image does not contain substation equipment when the similarity is less than the similarity threshold.
[0163] The substation equipment identification device provided in the embodiment of the present application can be applied to substation equipment identification devices, such as PC terminals, cloud platforms, servers and server clusters. Figure 3 The hardware structure diagram of the substation equipment identification device is shown. Figure 3 ,The hardware structure of the substation equipment identification device may include: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;
[0164] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0165] The processor 1 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention;
[0166] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory;
[0167] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to:
[0168] Obtain substation images and an improved comparative language-image pre-trained CLIP model. In the improved CLIP model, an adapter module is added to each Vision Transformer module in the vision network, and the language network is replaced with a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation devices.
[0169] The improved CLIP model is used to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolutional layer is used to process the image features to determine the equipment identification of the substation equipment area; if not, a reminder is generated to indicate that the substation image does not contain substation equipment.
[0170] Optionally, the refined functions and extended functions of the program may refer to the above description.
[0171] The present application also provides a readable storage medium, which may store a program suitable for execution by a processor, wherein the program is used to:
[0172] Obtain substation images and an improved comparative language-image pre-trained CLIP model. In the improved CLIP model, an adapter module is added to each Vision Transformer module in the vision network, and the language network is replaced with a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation devices.
[0173] The improved CLIP model is used to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolutional layer is used to process the image features to determine the equipment identification of the substation equipment area; if not, a reminder is generated to indicate that the substation image does not contain substation equipment.
[0174] Optionally, the refined functions and extended functions of the program may refer to the above description.
[0175] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0176] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0177] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. The various embodiments of the present application may be combined with each other. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying substation equipment, characterized in that: include: Obtain substation images and an improved comparative language-image pre-trained CLIP model. In the improved CLIP model, an adapter module is added to each Vision Transformer module in the vision network, and the language network is replaced with a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation devices. The improved CLIP model is used to identify whether there is a substation equipment area in the substation image. If so, the visual network is used to generate image features corresponding to the substation equipment area, and the language encoding convolutional layer is used to process the image features to determine the equipment identification of the substation equipment area; if not, a reminder is generated to indicate that the substation image does not contain substation equipment.
2. The substation equipment identification method according to claim 1, characterized in that: Get the improved CLIP model, including: Build a CLIP model that includes a visual network and a language network; In the CLIP model, an Adapter module is added after the multi-head self-attention layer of each Vision Transformer module in the vision network. At least two MLP modules are added before the input layer of the language network in the CLIP model. Each MLP module consists of a three-layer linear layer mixed with a Layer Norm module with a residual branch to form the initial CLIP model. Acquire multiple training samples, each training sample comprising a training device image and a corresponding training device code, wherein the training device images have the same size; Using each training sample, the initial CLIP model is trained until the initial CLIP model meets a preset stopping condition; The language network of the final initial CLIP model is updated to a language encoding convolutional layer to generate an improved CLIP model.
3. The substation equipment identification method according to claim 2, characterized in that: The obtaining of multiple training samples includes: Obtain different training images marked with training device image areas and their corresponding training device identifiers, where all training device identifiers of each training image constitute identifiers corresponding to all substation devices; Extracting a training device image region from each training image and normalizing the size of each training device image region to obtain each training device image; Using the same coding template, based on each training device identifier, obtain the training device code corresponding to each training device identifier; Each training device code is used as the annotation label of the corresponding training device image to form a training sample.
4. The substation equipment identification method according to claim 3, characterized in that: The language network of the final initial CLIP model is updated to a language encoding convolutional layer to generate an improved CLIP model, including: In combination with the coding template, based on the identification of each substation equipment, a language description corresponding to each substation equipment is generated; Perform dimension processing and numerical scaling on each language description to generate a device language representation for each substation device; Initialize the language network as a convolutional layer with a kernel size of 1x1; The conv_weight unit in the convolutional layer is replaced with each device language representation, and the conv_bias unit is replaced with a logit_bias unit to form a language encoding convolutional layer, thereby obtaining an improved CLIP model.
5. The substation equipment identification method according to claim 2, characterized in that: The initial CLIP model is trained using the training samples, including: In each iterative training, each training sample is sequentially input into the initial CLIP model to obtain the prediction result of the initial CLIP model; Calculate the prediction loss value based on each training sample and its corresponding prediction result; Based on the predicted loss value, parameters of each Adapter module and each MLP module are adjusted.
6. The substation equipment identification method according to claim 5, characterized in that: The calculation of the prediction loss value based on each training sample and its corresponding prediction result includes: Obtain the image training features and text training features extracted by the initial CLIP model based on each training sample and calculate the similarity matrix; In combination with the Sigmoid loss function, the prediction loss value is calculated based on each training sample and its corresponding prediction result, and the similarity matrix.
7. The substation equipment identification method according to claim 1, characterized in that: After processing the image features using the language encoding convolution layer to determine the equipment identification of the substation equipment area, the method further includes: Based on the improved CLIP model, a device language representation corresponding to the device identifier is obtained, and a similarity between the corresponding device language representation and the image feature is calculated. When the similarity exceeds a similarity threshold, the device identifier is output; when the similarity is less than the similarity threshold, a step of generating a reminder indicating that the substation image does not contain substation equipment is executed.
8. A substation equipment identification device, characterized in that: include: An acquisition module, used to acquire substation images and an improved comparative language-image pre-trained CLIP model. In the improved CLIP model, an adapter module is added to each Vision Transformer module in the vision network, and the language network is replaced with a language encoding convolutional layer. The language encoding convolutional layer is injected with the device language representation corresponding to all substation devices; an identification module for using the improved CLIP model to identify whether a substation equipment area exists in the substation image; if so, generating image features corresponding to the substation equipment area using the visual network, and processing the image features using the language encoding convolutional layer to determine the equipment identification of the substation equipment area; if not, generating a reminder indicating that the substation image does not contain substation equipment.
9. A substation equipment identification device, characterized in that: including memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the substation equipment identification method according to any one of claims 1 to 7.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the substation equipment identification method according to any one of claims 1 to 7 is implemented.