A method, system, device and medium for detecting anomalies of a multi-modal large model of electric power equipment

By partially updating the target power scene based on the multimodal large model of general power scenes, the problems of low abnormal detection efficiency and insufficient accuracy in the power grid are solved, and more efficient and accurate abnormal detection of power equipment is achieved.

CN119206596BActive Publication Date: 2025-05-23STATE GRID INTELLIGENCE TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411719259.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-23
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The prior art has problems of low construction efficiency and insufficient accuracy in the abnormal detection of power equipment in the power grid. Especially in the new power grid and new energy access environment, traditional convolutional neural network models are difficult to meet the needs of massive data processing and diversified anomaly identification.

Method used

A multimodal large model anomaly detection method for power equipment is proposed. By obtaining the general anomaly detection visual sub-model, a general text encoder and a general language sub-model of general power scenes, and updating these sub-models based on the sample data of the target power scene, forming a multimodal large model of the target power scene to improve the accuracy and construction efficiency of abnormality detection.

Benefits of technology

By partially updating the model structure and parameters, it can better adapt to specific power scenarios, improve the accuracy and efficiency of abnormal detection, reduce computing resource consumption, and reduce storage requirements and difficulty in use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206596B_ABST
    Figure CN119206596B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for detecting anomalies of a multimodal large model of electric power equipment. The method comprises: obtaining a general anomaly detection visual submodel, a general text encoder and a general language submodel of a general electric power scene; obtaining a multimodal large model of electric power equipment of a target electric power scene based on the general anomaly detection visual submodel, the general text encoder and the general language submodel, as well as a second sample image of a target electric power scene, a second prompt text associated with the second sample image, a second anomaly detection question text associated with the second sample image, a second anomaly detection positioning label associated with the second sample image and a second anomaly detection answer label associated with the second sample image; using the multimodal large model of electric power equipment of the target electric power scene to perform anomaly detection on an image to be detected of the target electric power scene. The embodiment of the present invention can improve the construction efficiency of the multimodal large model of electric power equipment and the accuracy of anomaly detection of an image to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, system, device and medium for detecting anomalies of a multi-modal large model of electric power equipment. Background Art

[0002] With the continuous expansion of the scale of power grids and the significant increase in the proportion of new energy access, the demand for new power grids is also constantly upgrading, and with it the accumulation and precipitation of massive data, which makes the traditional manual inspection method face unprecedented pressure. Professional visual models built based on traditional convolutional neural networks have been widely used in power grid business. Compared with these models, large models have shown greater application potential in multiple professional fields such as safety supervision and equipment maintenance with their large number of parameters, higher accuracy and stronger generalization ability.

[0003] The power vision big model uses deep learning technology and computer vision methods to process various types of data related to the power system, aiming to promote the transformation of the power equipment operation and maintenance mode from traditional manual operation and maintenance to intelligent operation and maintenance. Summary of the invention

[0004] The present invention provides a method, system, device and medium for detecting anomaly of a multi-modal large model of electric power equipment, so as to improve the construction efficiency and accuracy of anomaly detection of a large model for anomaly detection in an electric power system.

[0005] According to one aspect of the present invention, a method for detecting anomalies of a multimodal large model of power equipment is provided, comprising: obtaining a general anomaly detection visual submodel, a general text encoder and a general language submodel of a general power scenario; wherein the general anomaly detection visual submodel, the general text encoder and the general language submodel are trained using a first sample image of a general power scenario, a first prompt text associated with the first sample image, a first anomaly detection question text, a first anomaly detection positioning label and a first anomaly detection answer label; inputting a second sample image of a target power scenario and a second prompt text associated with the second sample image into the general anomaly detection visual submodel to obtain a second anomaly detection positioning prediction result, and using the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update a general feature extraction adaptation network and a general detection positioning layer in the general anomaly detection visual submodel to obtain a target anomaly detection visual submodel; inputting the second sample image of the target power scenario and the second prompt text associated with the second sample image into the general anomaly detection visual submodel to obtain a second anomaly detection positioning prediction result, and using the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update a general feature extraction adaptation network and a general detection positioning layer in the general anomaly detection visual submodel to obtain a target anomaly detection visual submodel; The image and the second prompt text are input into the target abnormality detection visual sub-model to obtain the target fusion feature, and the second abnormality detection question text associated with the second sample image is input into the universal text encoder to obtain the universal question text feature, and the target fusion feature and the universal question text feature are fused to obtain the second feature fusion result; the second feature fusion result is input into the universal language sub-model to obtain the second abnormality detection answer prediction result, and the universal text encoder and the universal language sub-model are updated using the second abnormality detection answer prediction result and the second abnormality detection answer label associated with the second sample image to obtain the target text encoder and target language sub-model of the target power scene; wherein the multimodal large model of power equipment of the target power scene includes the target abnormality detection visual sub-model, the target text encoder and the target language sub-model; the multimodal large model of power equipment of the target power scene is used to perform abnormality detection on the image to be detected of the target power scene.

[0006] According to another aspect of the present invention, a multimodal large model anomaly detection system for power equipment is provided, comprising: a general model acquisition module, used to acquire a general anomaly detection visual submodel, a general text encoder and a general language submodel of a general power scenario; wherein the general anomaly detection visual submodel, the general text encoder and the general language submodel are trained using a first sample image of a general power scenario, a first prompt text associated with the first sample image, a first anomaly detection question text, a first anomaly detection positioning label and a first anomaly detection answer label; an anomaly detection visual module, used to input a second sample image of a target power scenario and a second prompt text associated with the second sample image into the general anomaly detection visual submodel to obtain a second anomaly detection positioning prediction result, and use the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update the general feature extraction adaptation network and the general detection positioning layer in the general anomaly detection visual submodel to obtain a target anomaly detection visual submodel; a feature fusion module, used to The second sample image and the second prompt text are input into the target anomaly detection visual sub-model to obtain a target fusion feature, and the second anomaly detection question text associated with the second sample image is input into the general text encoder to obtain a general question text feature, and the target fusion feature and the general question text feature are fused to obtain a second feature fusion result; a text language module is used to input the second feature fusion result into the general language sub-model to obtain a second anomaly detection answer prediction result, and use the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image to update the general text encoder and the general language sub-model to obtain a target text encoder and a target language sub-model for a target power scene; wherein the multimodal large model of power equipment for the target power scene includes the target anomaly detection visual sub-model, the target text encoder and the target language sub-model; a prediction module is used to use the multimodal large model of power equipment for the target power scene to perform anomaly detection on the image to be detected of the target power scene.

[0007] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the multi-modal large model anomaly detection method for power equipment according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-modal large model anomaly detection method for power equipment described in any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multi-modal large model anomaly detection method for power equipment described in any embodiment of the present invention when executed.

[0010] The embodiment of the present invention updates some submodules in the general anomaly detection visual submodel based on the sample data of the target power scene, thereby obtaining a target anomaly detection visual submodel suitable for the target power scene, while retaining the generalization ability of the target anomaly detection visual submodel for other scenes, so that the target anomaly detection visual submodel can also better adapt to the target power scene, and realize image positioning prediction in the corresponding specific scene. The target text encoder and target language submodel are updated according to the sample data of the target power scene, thereby obtaining a target text encoder and target language submodel for the target power scene, while retaining the generalization ability of the target text encoder and target language submodel for other scenes, so that the target text encoder and target language submodel can also better adapt to the target power scene, and realize anomaly detection answer prediction in the corresponding specific scene. By introducing sample data of the target power scenario, the multimodal large model of power equipment in the general power scenario is optimized to obtain the multimodal large model of power equipment in the target power scenario, so that the multimodal large model of power equipment in the target power scenario can learn the unique characteristics and laws of the target power scenario, thereby improving the accuracy of anomaly detection of the multimodal large model of power equipment in the target power scenario in the corresponding power scenario. Based on the multimodal large model of power equipment in the general power scenario, some sub-modules are updated, and the consumption of computing resources can be reduced by optimizing the local model structure and some parameters. The multimodal large models of power equipment corresponding to different specific scenarios have the same structure, but the parameters on some sub-modules are different, so that accurate anomaly detection in different specific scenarios can be achieved, without deploying multiple models in the system, reducing storage requirements and difficulty of use. Phased training helps to reduce the number of parameters that need to be updated in each training, thereby reducing the consumption of computing resources and speeding up training.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0013] Figure 1 This is a first flow chart of a method for detecting anomalies of a multi-modal large model of electric power equipment provided by an embodiment of the present invention.

[0014] Figure 2 This is a second flow chart of a method for detecting anomalies of a multi-modal large model of electric power equipment provided in an embodiment of the present invention.

[0015] Figure 3 It is a schematic diagram of the model structure of a large multimodal model of power equipment in a general power scenario.

[0016] Figure 4 It is a schematic diagram of the structure of the deformable attention layer.

[0017] Figure 5 It is a schematic diagram of the model structure of the multimodal large model of power equipment in the target power scenario.

[0018] Figure 6 It is a structural schematic diagram of a multi-modal large model anomaly detection system for power equipment provided by an embodiment of the present invention.

[0019] Figure 7 It is a schematic diagram of the structure of an electronic device implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] Figure 1 This is the first flow chart of a method for detecting anomalies of a large multi-modal model of electric power equipment provided by an embodiment of the present invention. This embodiment is applicable to the case of detecting anomalies of an image to be detected in a target electric power scene. The method can be executed by a large multi-modal model anomaly detection system for electric power equipment. The system can be implemented in the form of hardware and / or software. The system can be configured in an electronic device with corresponding data processing capabilities. Figure 1 As shown, the method includes.

[0023] S110, obtaining a universal anomaly detection visual sub-model, a universal text encoder, and a universal language sub-model for universal power scenarios.

[0024] Among them, the general anomaly detection visual sub-model, the general text encoder and the general language sub-model are trained using the first sample image of a general power scene, the first prompt text associated with the first sample image, the first anomaly detection question text, the first anomaly detection location label and the first anomaly detection answer label.

[0025] General power scenarios can be subdivided into specific scenarios, including power transmission scenarios, substation scenarios, and power distribution scenarios. Collect abnormal equipment images in general power scenarios, such as damaged towers, broken wires, and transformer oil leakage. Use data enhancement methods such as geometric transformation, random cropping, scale change, and color perturbation to expand the abnormal equipment images to increase the diversity of samples and thus improve the generalization ability of the model. The expanded abnormal equipment images in general power scenarios are used as the first sample images.

[0026] A prompt text is set for the first sample image as the first prompt text associated with the first sample image. The first prompt text is used to provide key prompts of the equipment scene and equipment target in the first sample image, such as "This is an image of a substation, and the image contains transformer equipment". The first prompt text is input by the user. A first anomaly detection positioning label, a first anomaly detection question text, and a first anomaly detection answer label are set for the first sample image. The first anomaly detection positioning label is the actual positioning result of the equipment target in the first sample image. The first anomaly detection question text is the question text configured for requesting to obtain the anomaly detection result of the first sample image, such as "Is there an abnormal area in this image?" or "What is the type of anomaly in this image?". The first anomaly detection answer label is the actual target anomaly detection result of the first sample image. At the same time, in order to roughly determine the location of the anomaly, the image is divided into 3×3 grids, named upper left, upper, upper right, left, middle, right, lower left, lower, and lower right, respectively, so that the first anomaly detection answer label can generate descriptive content according to the location of the anomaly, such as "Yes, there is damage in the middle area of ​​the input image."

[0027] The multimodal large model of power equipment in a general power scenario includes a general anomaly detection visual sub-model, a general text encoder and a general language sub-model. The general anomaly detection visual sub-model, the general text encoder and the general language sub-model are trained using a first sample image of a general power scenario, a first prompt text associated with the first sample image, a first anomaly detection question text, a first anomaly detection location label and a first anomaly detection answer label.

[0028] The number of power equipment is huge and varied, and equipment abnormalities are diverse. Anomaly detection in general power scenarios has the problem of poor generalization, and an independent anomaly detection model is constructed for each specific anomaly detection task in general power scenarios. Multiple models need to be deployed on multiple machines, which results in low detection efficiency and extremely high resource consumption. This application constructs a multimodal large model of power equipment for specific scenarios based on a multimodal large model of power equipment in general power scenarios, realizes rapid iteration of the multimodal large model of power equipment in specific scenarios, better adapts to the identification requirements of equipment anomalies in specific scenarios, and thus significantly reduces the development cycle and resource consumption.

[0029] S120. Input the second sample image of the target power scene and the second prompt text associated with the second sample image into the universal anomaly detection visual sub-model to obtain a second anomaly detection positioning prediction result, and use the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update the universal feature extraction adaptation network and the universal detection positioning layer in the universal anomaly detection visual sub-model to obtain the target anomaly detection visual sub-model.

[0030] Any specific scenario in the general power scenario is taken as a target power scenario, and a large multimodal model of power equipment of the target power scenario is constructed. The large multimodal model of power equipment of the target power scenario includes a target anomaly detection visual sub-model. The target anomaly detection visual sub-model is used to locate and predict the abnormal equipment in the second sample image according to the second sample image of the target power scenario and the second prompt text associated with the second sample image.

[0031] Specifically, collect and expand the equipment abnormality image of the target power scene as the second sample image, set the second prompt text associated with the second sample image, the second abnormality detection positioning label associated with the second sample image, and set the second abnormality detection question and the second abnormality detection answer label associated with the second sample image. On the basis of the general abnormality detection visual sub-model, according to the second sample image of the target power scene, the second prompt text associated with the second sample image, and the second abnormality detection positioning label associated with the second sample image, some sub-modules in the general abnormality detection visual sub-model are updated to obtain a target abnormality detection visual sub-model suitable for the target power scene. Among them, the general abnormality detection visual sub-model includes at least a general feature extraction adaptation network and a general detection positioning layer. The general feature extraction adaptation network and the general detection positioning layer in the general abnormality detection visual sub-model are updated by the second sample image of the target power scene, the second prompt text, and the second abnormality detection positioning label to obtain the target abnormality detection visual sub-model.

[0032] By constructing a target anomaly detection visual sub-model suitable for the target power scenario based on the general anomaly detection visual sub-model, the target anomaly detection visual sub-model can better adapt to the target power scenario while retaining the generalization ability of the target anomaly detection visual sub-model for other scenarios, and realize image positioning prediction in the corresponding specific scenarios.

[0033] S130. Input the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain a target fusion feature, and input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the target fusion feature and the general question text feature to obtain a second feature fusion result.

[0034] S140. Input the second feature fusion result into the general language sub-model to obtain a second anomaly detection answer prediction result, and use the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image to update the general text encoder and the general language sub-model to obtain a target text encoder and target language sub-model for the target power scenario.

[0035] Among them, the multimodal large model of power equipment in the target power scenario also includes a target text encoder and a target language sub-model. The target text encoder is used to extract the general question text features of the second anomaly detection question text associated with the second sample image. The second feature fusion result obtained by fusing the target fusion feature and the general question text feature is input into the general language sub-model, and the general language sub-model outputs the corresponding second anomaly detection answer prediction result based on the second feature fusion result. For example, if the second anomaly detection question text is "Is there an abnormal area in this picture", the corresponding second anomaly detection answer prediction result can be "Yes, there is an abnormal area in the upper left corner of this image"; for another example, if the second anomaly detection question text is "What is the type of anomaly in this picture", the corresponding second anomaly detection answer prediction result can be "There is oil leakage on the surface of the transformer in the picture".

[0036] The target fusion feature is obtained by inputting the second sample image and the second prompt text into the target anomaly detection visual sub-model. The target fusion feature includes not only the image feature of the second sample image but also the text feature of the second prompt text, thereby realizing multimodal feature fusion, which is beneficial for coping with complex scenes and improving the accuracy of image positioning prediction. The general question text feature is obtained by inputting the second anomaly detection question text associated with the second sample image into the general text encoder, and the target fusion feature and the general question text feature are fused to obtain the second feature fusion result, and the second feature fusion result is input into the general language sub-model to obtain the second anomaly detection answer prediction result. The general language sub-model can obtain the second anomaly detection answer prediction result of the second sample image based on the target fusion feature under the guidance of the general question text feature, so that the user can quickly determine the anomaly-related information in the second sample image according to the second anomaly detection answer prediction result. By updating the general text encoder and the general language sub-model according to the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image, the target text encoder and target language sub-model of the target power scenario are obtained. While retaining the generalization ability of the target text encoder and the target language sub-model for other scenarios, the target text encoder and the target language sub-model can also better adapt to the target power scenario and realize the anomaly detection answer prediction in the corresponding specific scenario.

[0037] Among them, the general language sub-model and the target language sub-model can be a large language model (Large Language Model, LLM), which is pre-trained based on a large amount of question and answer data from various fields on the Internet.

[0038] By decomposing the large multimodal model of power equipment in the target power scenario into relatively independent target anomaly detection visual sub-models, target text encoders, and target language sub-models, each sub-module can be optimized relatively independently, and the most appropriate training strategy, loss function, and evaluation index can be selected according to the characteristics of each sub-module. Since each sub-module only focuses on its specific task during training, it can converge to the optimal solution faster, thereby improving the training effect and model performance. Different sub-modules can be combined and replaced according to actual needs to adapt to different application scenarios and needs, improving the application flexibility of the large multimodal model of power equipment.

[0039] The multimodal large model of power equipment of the target power scenario is trained in stages. In the first stage, the second sample image in the sample data of the target power scenario and the second prompt text associated with the second sample image are input into the general anomaly detection visual sub-model to obtain the second anomaly detection positioning prediction result, and the general anomaly detection visual sub-model is updated according to the second anomaly detection positioning prediction result and the second anomaly detection positioning label of the second sample image to obtain the target anomaly detection visual sub-model. In the second stage, the target fusion feature is obtained by inputting the second sample image and the second prompt text into the target anomaly detection visual sub-model; the general question text feature is obtained by inputting the second anomaly detection question text associated with the second sample image into the general text encoder, and the target fusion feature and the general question text feature are fused to obtain the second feature fusion result. The second feature fusion result is used as the input of the general language sub-model to obtain the second anomaly detection answer prediction result, and the general text encoder and the general language sub-model are updated according to the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image to obtain the target text encoder and the target language sub-model. That is, based on the training completed in the first stage, the second stage uses the output of the target anomaly detection visual sub-model obtained in the first stage to train the general text encoder and the general language sub-model to obtain the target text encoder and target language sub-model for the target power scenario.

[0040] Dividing the training process of the multimodal large model of power equipment in the target power scenario into the training process of the target anomaly detection visual sub-model and the training process of the target text encoder and target language sub-model, and conducting phased training helps to reduce the number of parameters that need to be updated in each training, thereby reducing the consumption of computing resources and speeding up training.

[0041] S150: Use a large multimodal model of electric power equipment in the target electric power scene to perform anomaly detection on the image to be detected in the target electric power scene.

[0042] The embodiment of the present invention updates some submodules in the general anomaly detection visual submodel based on the sample data of the target power scene, thereby obtaining a target anomaly detection visual submodel suitable for the target power scene, while retaining the generalization ability of the target anomaly detection visual submodel for other scenes, so that the target anomaly detection visual submodel can also better adapt to the target power scene, and realize image positioning prediction in the corresponding specific scene. The target text encoder and target language submodel are updated according to the sample data of the target power scene, thereby obtaining a target text encoder and target language submodel for the target power scene, while retaining the generalization ability of the target text encoder and target language submodel for other scenes, so that the target text encoder and target language submodel can also better adapt to the target power scene, and realize anomaly detection answer prediction in the corresponding specific scene. By introducing sample data of the target power scenario, the multimodal large model of power equipment in the general power scenario is optimized to obtain the multimodal large model of power equipment in the target power scenario, so that the multimodal large model of power equipment in the target power scenario can learn the unique characteristics and laws of the target power scenario, thereby improving the accuracy of anomaly detection of the multimodal large model of power equipment in the target power scenario in the corresponding power scenario. Based on the multimodal large model of power equipment in the general power scenario, some sub-modules are updated, and the consumption of computing resources can be reduced by optimizing the local model structure and some parameters. The multimodal large models of power equipment corresponding to different specific scenarios have the same structure, but the parameters on some sub-modules are different, which can realize accurate anomaly detection in the corresponding specific scenarios, without deploying multiple models in the system, reducing storage requirements and difficulty of use. Phased training helps to reduce the number of parameters that need to be updated in each training, thereby reducing the consumption of computing resources and speeding up training.

[0043] In an optional embodiment, the general anomaly detection visual submodel, the general text encoder and the general language submodel are trained by the following operations: a first sample image of a general power scene and a first prompt text associated with the first sample image are input into an initial anomaly detection visual submodel to obtain a first anomaly detection positioning prediction result, and the initial anomaly detection visual submodel is updated using the first anomaly detection positioning prediction result and the first anomaly detection positioning label of the first sample image to obtain the general anomaly detection visual submodel; the first sample image and the first prompt text are input into the general anomaly detection visual submodel to obtain a general fusion feature, and the first anomaly detection question text associated with the first sample image is input into the initial text encoder to obtain an initial question text feature, and the general fusion feature and the initial question text feature are fused to obtain a first feature fusion result; the first feature fusion result is input into the initial language submodel to obtain a first anomaly detection answer prediction result, and the initial text encoder and the initial language submodel are updated using the first anomaly detection answer prediction result and the first anomaly detection answer label associated with the first sample image to obtain the general text encoder and the general language submodel.

[0044] The multimodal large model of power equipment in the general power scenario is trained in stages. In the first stage, the first sample image and the first prompt text associated with the first sample image in the sample data of the general power scenario are input into the initial anomaly detection visual sub-model to obtain the first anomaly detection positioning prediction result, and the initial anomaly detection visual sub-model is updated according to the first anomaly detection positioning prediction result and the first anomaly detection positioning label of the first sample image to obtain the general anomaly detection visual sub-model. In the second stage, the first sample image and the first prompt text are input into the general anomaly detection visual sub-model to obtain the general fusion feature; the first anomaly detection question text associated with the first sample image is input into the initial text encoder to obtain the initial question text feature, and the general fusion feature and the initial question text feature are fused to obtain the first feature fusion result. The first feature fusion result is used as the input of the initial language sub-model to obtain the first anomaly detection answer prediction result, and the first anomaly detection answer prediction result and the first anomaly detection answer label associated with the first sample image are used to update the initial text encoder and the initial language sub-model to obtain the general text encoder and the general language sub-model. That is, based on the training completed in the first stage, the second stage uses the output of the general anomaly detection visual sub-model obtained in the first stage to train the initial text encoder and the initial language sub-model to obtain the general text encoder and general language sub-model for general power scenarios.

[0045] Dividing the training process of the multimodal large model of power equipment in general power scenarios into the training process of the general anomaly detection visual sub-model and the training process of the general text encoder and general language sub-model, and conducting phased training helps reduce the number of parameters that need to be updated during each training, thereby reducing the consumption of computing resources and speeding up training.

[0046] Figure 2 This is a second flow chart of a method for detecting abnormalities in a multi-modal large model of a power device provided by an embodiment of the present invention. This embodiment is optimized and improved on the basis of the above embodiment. Figure 2 As shown, the method includes.

[0047] S210, obtaining a general anomaly detection visual sub-model, a general text encoder, and a general language sub-model for a general power scenario.

[0048] Among them, the general anomaly detection visual sub-model, the general text encoder and the general language sub-model are trained using the first sample image of a general power scene, the first prompt text associated with the first sample image, the first anomaly detection question text, the first anomaly detection location label and the first anomaly detection answer label.

[0049] Among them, Figure 3 As shown, the universal anomaly detection visual sub-model includes a universal image feature extraction network, a universal feature extraction adaptation network, a universal text feature extraction network, a universal feature enhancement network and a universal detection and positioning layer.

[0050] S220, inputting a second sample image of the target power scene into the universal image feature extraction network, and processing the universal image feature extraction network and the universal feature extraction adaptation network to obtain a scale-adapted universal image feature.

[0051] The general image feature extraction network is based on the deep learning model architecture of the hierarchical visual self-attention mechanism using shifted windows (Hierarchical Vision Transformer using Shifted Windows, SwinTransformer), including an image feature segmentation module (Patch Partition) and 4 stages (stage). Stage 1 consists of a linear embedding module (Linear Embedding) and a Swin Transformer layer, and stages 2, 3, and 4 are all composed of a feature merging module (Patch Merging) and a Swin Transformer layer. The second sample image input is divided into image blocks by the image feature segmentation module, flattened in the channel dimension by the linear embedding module of stage 1, and then input into the SwinTransformer layer and stages 2 to 4 for feature extraction to obtain the initial image features. The feature merging module implements the downsampling of the initial image features to obtain the general image features. The Swin Transformer layer consists of a Layer Norm (normalization) layer, a Windows Multi-head Self-Attention (window multi-head attention mechanism) layer, a Shifted Windows Multi-head Self-Attention (shifted window multi-head self-attention) layer, a Layer Norm layer, and an MLP (MultilayerPerceptron) layer.

[0052] The shape of power equipment is usually irregular. In order to improve the accuracy of positioning and prediction of irregular abnormal targets of power equipment, a feature extraction adaptation network is proposed. The feature extraction adaptation network contains multiple deformable attention layers. By introducing offsets, the attention to the irregular abnormal areas of power equipment is strengthened. The structure of the deformable attention layer is shown in Figure 2. Figure 4 As shown. The deformable attention layer is mainly composed of 1×1 convolution, GELU (Gaussian Error Linear Unit) activation function, deformable convolution, and deep deformable dilated convolution. Its core is deformable convolution, which introduces an offset offset on the basis of traditional convolution, so that when the convolution kernel samples on the input feature map, the receptive field is deformed and is no longer a regular square, but closer to the target shape. Deep deformable dilated convolution adds a hyperparameter expansion rate d to the deformable convolution operation. For example, d can be taken as 2 to expand the receptive field during sampling without increasing the amount of calculation. Through the above operations, the receptive field is closer to the shape and size of the target of attention, thereby enhancing the boundaries of abnormal targets of the device.

[0053] The second sample image of the target power scene is input into the universal image feature extraction network to obtain multi-scale universal image features. The multi-scale universal image features are input into the universal feature extraction adaptation network to further obtain multi-scale universal image features adapted to multi-morphological targets.

[0054] S230: Input the second prompt text associated with the second sample image into the universal text feature extraction network, and obtain the re-parameterized universal text features through the universal text feature extraction network and the initial re-parameterization layer.

[0055] Both the general text feature extraction network and the target text feature extraction network can select a network based on BERT (Bidirectional Encoder Representations from Transformers) to extract text features of the prompt text. A reparameterization layer is added to the model structure of the general anomaly detection visual submodel for the target power scenario, which is used to learn the general anomaly detection visual submodel to a specific scenario based on sample data of a specific scenario, so that the general anomaly detection visual submodel with the added initial reparameterization layer can learn the specific scene information in the second prompt text associated with the second sample image through the general text feature extraction network and the initial reparameterization layer.

[0056] The second prompt text of the second sample image is input into the universal text feature extraction network to obtain the universal text feature of the second sample image. The universal text feature of the second sample image is input into the initial reparameterization layer, and the universal text feature is scaled by scaling and offset operations to enhance the original feature expression ability and obtain the reparameterized universal text feature. The specific implementation process of reparameterization is shown in the following formulas (1)-(3).

[0057] (1)

[0058] (2)

[0059] in, is the scaling factor, is the offset, x is the output of the image hint text in the text feature extraction network. t is the input of the previous feature layer of the reparameterized layer, that is, t is the input of the text feature extraction network, w and b are the weight of the text feature extraction network and the bias of the text feature extraction network respectively. * represents the convolution operation in the convolution layer or the multiplication operation in the MLP layer. Based on formula (2), formula (1) can also be expressed as formula (3), which is as follows.

[0060] (3)

[0061] S240, inputting the scale-adapted universal image features and the re-parameterized universal text features into the universal feature enhancement network for cross-modal feature fusion to obtain universal fusion features, and inputting the universal fusion features into the universal detection and positioning layer to obtain a second anomaly detection and positioning prediction result.

[0062] The feature enhancement network is used to fuse image features and text features. Specifically, you can choose a feature pyramid network based on FPN (Feature Pyramid Network) to fuse the multi-scale image features that have passed through the deformable attention layer, and then pass the fused multi-scale features through a linear layer for feature alignment and text feature calculation. The detection and positioning layer is used to predict the abnormal location information of the device in the image.

[0063] The multi-scale universal image features adapted to multi-morphological targets and the re-parameterized universal text features are input into the universal feature enhancement network for cross-modal feature fusion to obtain universal fusion features, and the universal fusion features are input into the universal detection and positioning layer to obtain the second anomaly detection and positioning prediction results.

[0064] S250, using the second anomaly detection and positioning prediction result and the second anomaly detection and positioning label associated with the second sample image to update the general feature extraction adaptation network, the initial reparameterization layer and the general detection and positioning layer to obtain a target feature extraction adaptation network, a target reparameterization layer and a target detection and positioning layer.

[0065] Freeze the general image feature extraction network, the general text feature extraction network and the general feature enhancement network, that is, keep the parameters of the general image feature extraction network, the general text feature extraction network and the general feature enhancement network unchanged. Use the second anomaly detection and positioning prediction result and the second anomaly detection and positioning label associated with the second sample image to update only the general feature extraction adaptation network, the initial reparameterization layer and the general detection and positioning layer to obtain the target feature extraction adaptation network, the target reparameterization layer and the target detection and positioning layer. Through the specified fine-tuning method, update the network parameters of the deformable attention layer in the general feature extraction adaptation network to obtain the target feature extraction adaptation network to adapt to the abnormal targets in the specific scene. When the general text feature extraction network is frozen, refer to formula (3), that is, w and b are frozen. Since w and b are frozen, in the update process of the reparameterization layer The sum between them is updated, and the reparameterized layer is updated by the scaling factor and offset The reparameterization layer can reparameterize the output of the text feature extraction network according to the specific power scenario to enhance the feature expression and strengthen the specific scenario information in the text features.

[0066] like Figure 5 As shown, the target anomaly detection visual sub-model includes a general image feature extraction network, the target feature extraction adaptation network, the general text feature extraction network, the target reparameterization layer, the general feature enhancement network and the target detection positioning layer.

[0067] In an optional embodiment, the target reparameterization layer is used to perform a linear transformation on the output of the general text feature extraction network; the target feature extraction adaptation network includes multiple deformable attention layers for processing irregular abnormal areas of power equipment.

[0068] Specifically, the target feature extraction adaptation network has the same network structure as the general feature extraction adaptation network. The target feature extraction adaptation network contains multiple deformable attention layers, which introduces offsets to strengthen the attention to irregular and abnormal areas of power equipment. Referring to formula (1), the target parameterization layer is scaled by the factor and offset Perform a linear transformation on the output x of the general text feature extraction network.

[0069] The target feature extraction adaptation network makes the receptive field closer to the shape and size of the target of interest, thereby enhancing the boundaries of the device abnormal target, so that the target abnormality detection visual sub-model can optimize the image features so that the image features can accurately express the abnormal targets of different devices, improving the accuracy of abnormality detection positioning prediction. The reparameterization layer transforms the output of the general text feature extraction network, that is, the general text feature, into a completely linear transformation, so that the anomaly detection visual sub-model can learn the specific scene information in the prompt text of the image without adding any additional parameters and computational costs, and strengthen the specific scene information in the target text features, so that the target anomaly detection visual sub-model can adapt to the corresponding specific scene.

[0070] S260. Input the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain a target fusion feature, and input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the target fusion feature and the general question text feature to obtain a second feature fusion result.

[0071] In an optional implementation, the step of inputting the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain target fusion features, inputting the second anomaly detection question text associated with the second sample image into the universal text encoder to obtain universal question text features, and fusing the target fusion features and the universal question text features to obtain a second feature fusion result includes: inputting the second sample image into the universal image feature extraction network, obtaining scale-adapted target image features through the universal image feature extraction network and the target feature extraction adaptation network; inputting the second prompt text into the universal text feature extraction network, obtaining re-parameterized target text features through the universal text feature extraction network and the target re-parameterization layer; inputting the scale-adapted target image features and the re-parameterized target text features into the universal feature enhancement network for cross-modal feature fusion to obtain target fusion features; inputting the second anomaly detection question text associated with the second sample image into the universal text encoder to obtain universal question text features, and fusing the universal question text features and the target fusion features to obtain a second feature fusion result.

[0072] S270. Input the second feature fusion result into the general language sub-model to obtain a second anomaly detection answer prediction result, and use the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image to update the general text encoder and the general language sub-model to obtain a target text encoder and target language sub-model for the target power scenario.

[0073] Among them, the multimodal large model of power equipment in the target power scenario includes the target anomaly detection visual sub-model, the target text encoder and the target language sub-model.

[0074] Specifically, the multimodal large model of power equipment in the target power scenario is trained locally in stages. In the first stage, the general anomaly detection visual sub-model with an initial reparameterization layer is trained, and the general feature extraction adaptation network, the initial reparameterization layer and the general detection and positioning layer are updated. After the first stage of training is completed, the target feature extraction adaptation network, the target reparameterization layer and the target detection and positioning layer are obtained, thereby obtaining the target anomaly detection visual sub-model.

[0075] According to the second sample image of the target power scene and the second prompt text associated with the second sample image, a target fusion feature is obtained based on the target anomaly detection visual sub-model. The second anomaly detection question text associated with the second sample image is input into the general text encoder to obtain a general question text feature, and the general question text feature and the target fusion feature are fused to obtain a second feature fusion result.

[0076] In the second stage, the second feature fusion result is used as the input of the general language sub-model to obtain the second anomaly detection answer prediction result, and the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image are used to update the general text encoder and the general language sub-model to obtain the target text encoder and target language sub-model for the target power scene. That is, in the second stage, based on the training completed in the first stage, the output of the target anomaly detection visual sub-model obtained by the training in the first stage is used to train the general text encoder and the general language sub-model to obtain the target text encoder and target language sub-model for the target power scene.

[0077] Phased targeted training allows the model to focus on different learning tasks at different stages. During the phased training process, by freezing the trained target anomaly detection visual sub-model, the number of parameters that need to be updated during each training can be reduced, thereby reducing the consumption of computing resources and speeding up training. For large multimodal models of power equipment, phased training helps the model better integrate information from different modalities; by freezing and updating the network parameters corresponding to different modalities at different stages, it can ensure that the model fully utilizes the complementarity of multimodal data during training and improves the accuracy and robustness of anomaly detection.

[0078] S280: Use a large multimodal model of electric power equipment in the target electric power scene to perform anomaly detection on the image to be detected in the target electric power scene.

[0079] A method for fine-tuning a large multimodal model of power equipment that combines specified and re-parameterized methods is proposed to construct an equipment anomaly detection dataset for a specific target power scenario. The specified and re-parameterized fine-tuning methods are combined to update some sub-modules in the multimodal large model of power equipment for general power scenarios to obtain a large multimodal model of power equipment for a specific power scenario. This reduces the number of parameters that need to be updated during the training of the large model and realizes the rapid construction of a large multimodal model of power equipment.

[0080] A target feature extraction and adaptation technology for a large electric power vision model is proposed. A target feature extraction and adaptation network for electric power equipment is designed and implemented to learn the unique features and rules of the target electric power scene, significantly improving the detection accuracy of irregular and abnormal equipment targets in the target electric power scene.

[0081] The embodiment of the present invention sets multiple deformable attention layers in the general feature extraction adaptation network and the target feature extraction adaptation network, and introduces an offset to strengthen the attention to the irregular abnormal areas of the power equipment, so that the receptive field is closer to the shape and size of the target of attention, thereby enhancing the boundary of the abnormal target of the equipment, so that the abnormal detection visual sub-model can optimize the image features, so that the image features can accurately express the abnormal targets of different equipment, thereby improving the accuracy of abnormal detection positioning prediction.

[0082] The output of the general text feature extraction network, namely the general text feature, is linearly transformed through the reparameterization layer, so that the target anomaly detection visual sub-model can learn the specific scene information in the prompt text of the image without adding any additional parameters and computational costs, thereby strengthening the specific scene information in the target text feature, so that the target anomaly detection visual sub-model can adapt to the corresponding specific scene.

[0083] Through phased training, the large multimodal model of power equipment can focus on different learning tasks at different stages. In the phased training process, by freezing the trained target anomaly detection visual sub-model, the number of parameters that need to be updated during each training can be reduced, thereby reducing the consumption of computing resources and speeding up training. Phased training helps the large multimodal model of power equipment to better integrate information from different modalities; by freezing and updating the network parameters corresponding to different modalities at different stages, it can ensure that the model fully utilizes the complementarity of multimodal data during training and improves the accuracy and robustness of anomaly detection.

[0084] On the basis of the large multimodal model of power equipment in the general power scenario, a large multimodal model of power equipment in the target power scenario is constructed. By partially freezing the parameters of the large multimodal model of power equipment in the general power scenario, the risk of overfitting of the model during training can be reduced, and the knowledge and information in the general power scenario contained in these parameters can be retained, which helps the model to generalize better when facing new tasks or new data sets. The model can more easily adapt to tasks and data sets in different scenarios and achieve a wider range of transfer learning effects; improve the accuracy of anomaly detection of images to be detected in the target power scenario. Locally freezing parameters can make full use of the model parameters that have been trained in the general power scenario, and only update the parameters related to the specific scenario, which accelerates the training process of the large multimodal model of power equipment in the target power scenario and improves performance.

[0085] Figure 6 Schematic diagram of a multi-modal large model anomaly detection system for power equipment provided by an embodiment of the present invention. Figure 6 As shown, the system includes.

[0086] The general model acquisition module 310 is used to obtain a general anomaly detection visual sub-model, a general text encoder and a general language sub-model of a general power scenario; wherein the general anomaly detection visual sub-model, the general text encoder and the general language sub-model are trained using a first sample image of a general power scenario, a first prompt text associated with the first sample image, a first anomaly detection question text, a first anomaly detection location label and a first anomaly detection answer label.

[0087] The anomaly detection visual module 320 is used to input the second sample image of the target power scene and the second prompt text associated with the second sample image into the universal anomaly detection visual sub-model to obtain a second anomaly detection positioning prediction result, and use the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update the universal feature extraction adaptation network and the universal detection positioning layer in the universal anomaly detection visual sub-model to obtain a target anomaly detection visual sub-model.

[0088] The feature fusion module 330 is used to input the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain a target fusion feature, and input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the target fusion feature and the general question text feature to obtain a second feature fusion result.

[0089] The text language module 340 is used to input the second feature fusion result into the general language sub-model to obtain a second anomaly detection answer prediction result, and use the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image to update the general text encoder and the general language sub-model to obtain a target text encoder and a target language sub-model for the target power scenario; wherein the multimodal large model of power equipment in the target power scenario includes the target anomaly detection visual sub-model, the target text encoder and the target language sub-model.

[0090] The prediction module 350 is used to perform anomaly detection on the image to be detected of the target power scene by using the multi-modal large model of the power equipment of the target power scene.

[0091] The multi-modal large model anomaly detection system for electric power equipment provided in the embodiment of the present invention can execute the multi-modal large model anomaly detection method for electric power equipment provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0092] Optionally, the universal anomaly detection visual sub-model includes a universal image feature extraction network, a universal feature extraction adaptation network, a universal text feature extraction network, a universal feature enhancement network and a universal detection and positioning layer.

[0093] The anomaly detection visual module includes.

[0094] The universal image feature unit is used to input the second sample image of the target power scene into the universal image feature extraction network, and obtain the scale-adapted universal image features through processing by the universal image feature extraction network and the universal feature extraction adaptation network.

[0095] The universal text feature unit is used to input the second prompt text associated with the second sample image into the universal text feature extraction network, and obtain the re-parameterized universal text features through the universal text feature extraction network and the initial re-parameterization layer.

[0096] The anomaly detection and positioning unit is used to input the scale-adapted universal image features and the re-parameterized universal text features into the universal feature enhancement network to perform cross-modal feature fusion to obtain universal fusion features, and input the universal fusion features into the universal detection and positioning layer to obtain a second anomaly detection and positioning prediction result.

[0097] A visual module updating unit is used to update the general feature extraction adaptation network, the initial reparameterization layer and the general detection positioning layer using the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to obtain a target feature extraction adaptation network, a target reparameterization layer and a target detection positioning layer; the target anomaly detection visual submodel includes a general image feature extraction network, the target feature extraction adaptation network, the general text feature extraction network, the target reparameterization layer, the general feature enhancement network and the target detection positioning layer.

[0098] Optionally, the target reparameterization layer is used to perform a linear transformation on the output of the general text feature extraction network; the target feature extraction adaptation network includes multiple deformable attention layers for processing irregular abnormal areas of power equipment.

[0099] Optionally, the feature fusion module includes.

[0100] The target image feature unit is used to input the second sample image into the general image feature extraction network, and obtain the scale-adapted target image features through the general image feature extraction network and the target feature extraction adaptation network.

[0101] The target text feature unit is used to input the second prompt text into the universal text feature extraction network, and obtain the re-parameterized target text feature through processing by the universal text feature extraction network and the target re-parameterization layer.

[0102] The target fusion feature unit is used to input the scale-adapted target image features and the re-parameterized target text features into the universal feature enhancement network to perform cross-modal feature fusion to obtain target fusion features.

[0103] The second feature fusion unit is used to input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the general question text feature with the target fusion feature to obtain a second feature fusion result.

[0104] Optionally, the system also includes a general model training module, which is used to train the general anomaly detection visual sub-model, the general text encoder and the general language sub-model through the following operations: inputting the first sample image of the general power scene and the first prompt text associated with the first sample image into the initial anomaly detection visual sub-model to obtain a first anomaly detection positioning prediction result, and using the first anomaly detection positioning prediction result and the first anomaly detection positioning label of the first sample image to update the initial anomaly detection visual sub-model to obtain the general anomaly detection visual sub-model; inputting the first sample image and the first prompt text into the general anomaly detection visual sub-model to obtain a general fusion feature, and inputting the first anomaly detection question text associated with the first sample image into the initial text encoder to obtain an initial question text feature, and fusing the general fusion feature with the initial question text feature to obtain a first feature fusion result; inputting the first feature fusion result into the initial language sub-model to obtain a first anomaly detection answer prediction result, and using the first anomaly detection answer prediction result and the first anomaly detection answer label associated with the first sample image to update the initial text encoder and the initial language sub-model to obtain the general text encoder and the general language sub-model.

[0105] Optionally, the first anomaly detection positioning label is the actual positioning result of the device target in the first sample image; the first anomaly detection question text is the question text configured for requesting to obtain the anomaly detection result of the first sample image; and the first anomaly detection answer label is the actual target anomaly detection result of the first sample image, which is used to roughly determine the location of the anomaly.

[0106] The further described multi-modal large model anomaly detection system for electric power equipment can also execute the multi-modal large model anomaly detection method for electric power equipment provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0107] Figure 7 A schematic diagram of an electronic device 40 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0108] like Figure 7 As shown, the electronic device 40 includes at least one processor 41, and a memory connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 to the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0109] A number of components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0110] The processor 41 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 41 executes the various methods and processes described above, such as the multi-modal large model anomaly detection method for power equipment.

[0111] In some embodiments, the multimodal large model anomaly detection method for power equipment may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the multimodal large model anomaly detection method for power equipment described above may be performed. Alternatively, in other embodiments, the processor 41 may be configured to execute the multimodal large model anomaly detection method for power equipment in any other appropriate manner (e.g., by means of firmware).

[0112] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0114] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0115] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0116] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0117] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0118] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0119] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-modal large model anomaly detection method for power equipment, characterized in that: include: Obtain a universal anomaly detection visual submodel, a universal text encoder, and a universal language submodel of a universal power scenario; wherein the universal anomaly detection visual submodel, the universal text encoder, and the universal language submodel are trained using a first sample image of a universal power scenario, a first prompt text associated with the first sample image, a first anomaly detection question text, a first anomaly detection location label, and a first anomaly detection answer label; Inputting a second sample image of the target power scene and a second prompt text associated with the second sample image into the universal anomaly detection visual sub-model to obtain a second anomaly detection positioning prediction result, and using the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update the universal feature extraction adaptation network and the universal detection positioning layer in the universal anomaly detection visual sub-model to obtain a target anomaly detection visual sub-model; Input the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain a target fusion feature, input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the target fusion feature and the general question text feature to obtain a second feature fusion result; The second feature fusion result is input into the general language sub-model to obtain a second anomaly detection answer prediction result, and the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image are used to update the general text encoder and the general language sub-model to obtain a target text encoder and a target language sub-model of the target power scenario; wherein the multimodal large model of power equipment of the target power scenario includes the target anomaly detection visual sub-model, the target text encoder and the target language sub-model; The multimodal large model of the power equipment of the target power scene is used to perform anomaly detection on the image to be detected of the target power scene; Among them, the general power scenario can be subdivided into specific scenarios, including transmission scenarios, substation scenarios and distribution scenarios; the target power scenario is any specific scene in the general power scenario; the first sample image is obtained by collecting abnormal equipment images in the general power scenario and expanding the abnormal equipment images in the general power scenario through a data enhancement method; the second sample image is obtained by collecting and expanding abnormal equipment images in the target power scenario.

2. The method according to claim 1, characterized in that The general anomaly detection visual sub-model includes a general image feature extraction network, a general feature extraction adaptation network, a general text feature extraction network, a general feature enhancement network and a general detection and positioning layer; The second sample image of the target power scene and the second prompt text associated with the second sample image are input into the universal anomaly detection visual sub-model to obtain a second anomaly detection positioning prediction result, and the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image are used to update the universal feature extraction adaptation network and the universal detection positioning layer in the universal anomaly detection visual sub-model to obtain the target anomaly detection visual sub-model, including: Inputting a second sample image of the target power scene into the universal image feature extraction network, and processing the universal image feature extraction network and the universal feature extraction adaptation network to obtain a scale-adapted universal image feature; Inputting the second prompt text associated with the second sample image into the universal text feature extraction network, and obtaining the re-parameterized universal text features through the universal text feature extraction network and the initial re-parameterization layer; Input the scale-adapted universal image features and the re-parameterized universal text features into the universal feature enhancement network to perform cross-modal feature fusion to obtain universal fusion features, and input the universal fusion features into the universal detection and positioning layer to obtain a second anomaly detection and positioning prediction result; The second anomaly detection and positioning prediction result and the second anomaly detection and positioning label associated with the second sample image are used to update the general feature extraction adaptation network, the initial reparameterization layer and the general detection and positioning layer to obtain a target feature extraction adaptation network, a target reparameterization layer and a target detection and positioning layer; the target anomaly detection visual submodel includes a general image feature extraction network, the target feature extraction adaptation network, the general text feature extraction network, the target reparameterization layer, the general feature enhancement network and the target detection and positioning layer.

3. The method according to claim 2, characterized in that The target reparameterization layer is used to perform a linear transformation on the output of the general text feature extraction network; the target feature extraction adaptation network includes a plurality of deformable attention layers for processing irregular abnormal areas of power equipment.

4. The method according to claim 2, characterized in that: The step of inputting the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain a target fusion feature, inputting the second anomaly detection question text associated with the second sample image into the universal text encoder to obtain a universal question text feature, and fusing the target fusion feature with the universal question text feature to obtain a second feature fusion result includes: Inputting the second sample image into the general image feature extraction network, and obtaining the scale-adapted target image features through the general image feature extraction network and the target feature extraction adaptation network; Inputting the second prompt text into the universal text feature extraction network, and processing the universal text feature extraction network and the target reparameterization layer to obtain reparameterized target text features; Inputting the scale-adapted target image features and the re-parameterized target text features into the universal feature enhancement network to perform cross-modal feature fusion to obtain target fusion features; The second anomaly detection question text associated with the second sample image is input into the universal text encoder to obtain universal question text features, and the universal question text features are fused with the target fusion features to obtain a second feature fusion result.

5. The method according to claim 1, characterized in that The general anomaly detection visual sub-model, the general text encoder and the general language sub-model are obtained by training through the following operations: Inputting a first sample image of a general power scene and a first prompt text associated with the first sample image into an initial anomaly detection visual sub-model to obtain a first anomaly detection positioning prediction result, and using the first anomaly detection positioning prediction result and the first anomaly detection positioning label of the first sample image to update the initial anomaly detection visual sub-model to obtain the general anomaly detection visual sub-model; Inputting the first sample image and the first prompt text into the general anomaly detection visual sub-model to obtain a general fusion feature, inputting the first anomaly detection question text associated with the first sample image into the initial text encoder to obtain an initial question text feature, and fusing the general fusion feature with the initial question text feature to obtain a first feature fusion result; The first feature fusion result is input into the initial language sub-model to obtain a first anomaly detection answer prediction result, and the first anomaly detection answer prediction result and the first anomaly detection answer label associated with the first sample image are used to update the initial text encoder and the initial language sub-model to obtain the universal text encoder and the universal language sub-model.

6. The method according to claim 1, characterized in that The first anomaly detection positioning label is a real positioning result of the device target in the first sample image; the first anomaly detection question text is a question text configured for requesting to obtain the anomaly detection result of the first sample image; The first anomaly detection answer label is a true target anomaly detection result of the first sample image, which is used to roughly determine the location of the anomaly.

7. A multi-modal large model anomaly detection system for power equipment, characterized in that: The system comprises: A general model acquisition module, used to acquire a general anomaly detection visual sub-model, a general text encoder and a general language sub-model of a general power scenario; wherein the general anomaly detection visual sub-model, the general text encoder and the general language sub-model are trained using a first sample image of a general power scenario, a first prompt text associated with the first sample image, a first anomaly detection question text, a first anomaly detection location label and a first anomaly detection answer label; an anomaly detection visual module, used for inputting a second sample image of a target power scene and a second prompt text associated with the second sample image into a universal anomaly detection visual sub-model to obtain a second anomaly detection positioning prediction result, and using the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to update a universal feature extraction adaptation network and a universal detection positioning layer in the universal anomaly detection visual sub-model to obtain a target anomaly detection visual sub-model; a feature fusion module, configured to input the second sample image and the second prompt text into the target anomaly detection visual sub-model to obtain a target fusion feature, input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the target fusion feature with the general question text feature to obtain a second feature fusion result; a text language module, configured to input the second feature fusion result into the general language sub-model to obtain a second anomaly detection answer prediction result, and to update the general text encoder and the general language sub-model using the second anomaly detection answer prediction result and the second anomaly detection answer label associated with the second sample image, so as to obtain a target text encoder and a target language sub-model of a target power scenario; wherein the multimodal large model of power equipment of the target power scenario includes the target anomaly detection visual sub-model, the target text encoder and the target language sub-model; A prediction module, used to perform anomaly detection on the image to be detected of the target power scene using a large multimodal model of power equipment of the target power scene; Among them, the general power scenario can be subdivided into specific scenarios, including transmission scenarios, substation scenarios and distribution scenarios; the target power scenario is any specific scene in the general power scenario; the first sample image is obtained by collecting abnormal equipment images in the general power scenario and expanding the abnormal equipment images in the general power scenario through a data enhancement method; the second sample image is obtained by collecting and expanding abnormal equipment images in the target power scenario.

8. The system according to claim 7, characterized in that The general anomaly detection visual sub-model includes a general image feature extraction network, a general feature extraction adaptation network, a general text feature extraction network, a general feature enhancement network and a general detection and positioning layer; The anomaly detection visual module includes: a universal image feature unit, which is used to input a second sample image of the target power scene into the universal image feature extraction network, and obtain a universal image feature after scale adaptation through processing by the universal image feature extraction network and the universal feature extraction adaptation network; A universal text feature unit, used for inputting a second prompt text associated with a second sample image into the universal text feature extraction network, and obtaining a re-parameterized universal text feature through processing by the universal text feature extraction network and the initial re-parameterization layer; An anomaly detection and positioning unit, used for inputting the scale-adapted universal image features and the re-parameterized universal text features into the universal feature enhancement network to perform cross-modal feature fusion to obtain universal fusion features, and inputting the universal fusion features into the universal detection and positioning layer to obtain a second anomaly detection and positioning prediction result; A visual module updating unit is used to update the general feature extraction adaptation network, the initial reparameterization layer and the general detection positioning layer using the second anomaly detection positioning prediction result and the second anomaly detection positioning label associated with the second sample image to obtain a target feature extraction adaptation network, a target reparameterization layer and a target detection positioning layer; the target anomaly detection visual submodel includes a general image feature extraction network, the target feature extraction adaptation network, the general text feature extraction network, the target reparameterization layer, the general feature enhancement network and the target detection positioning layer.

9. The system according to claim 8, characterized in that The target reparameterization layer is used to perform a linear transformation on the output of the general text feature extraction network; the target feature extraction adaptation network includes a plurality of deformable attention layers for processing irregular abnormal areas of power equipment.

10. The system according to claim 8, characterized in that The feature fusion module comprises: A target image feature unit, used for inputting the second sample image into the general image feature extraction network, and obtaining the scale-adapted target image features through the general image feature extraction network and the target feature extraction adaptation network; A target text feature unit, used for inputting the second prompt text into the universal text feature extraction network, and obtaining the re-parameterized target text feature through processing by the universal text feature extraction network and the target re-parameterization layer; A target fusion feature unit, used for inputting the scale-adapted target image features and the re-parameterized target text features into the universal feature enhancement network to perform cross-modal feature fusion to obtain target fusion features; The second feature fusion unit is used to input the second anomaly detection question text associated with the second sample image into the general text encoder to obtain a general question text feature, and fuse the general question text feature with the target fusion feature to obtain a second feature fusion result.

11. The system according to claim 7, characterized in that It also includes a general model training module, which is used to train the general anomaly detection visual sub-model, the general text encoder and the general language sub-model through the following operations: Inputting a first sample image of a general power scene and a first prompt text associated with the first sample image into an initial anomaly detection visual sub-model to obtain a first anomaly detection positioning prediction result, and using the first anomaly detection positioning prediction result and the first anomaly detection positioning label of the first sample image to update the initial anomaly detection visual sub-model to obtain the general anomaly detection visual sub-model; Inputting the first sample image and the first prompt text into the general anomaly detection visual sub-model to obtain a general fusion feature, inputting the first anomaly detection question text associated with the first sample image into the initial text encoder to obtain an initial question text feature, and fusing the general fusion feature with the initial question text feature to obtain a first feature fusion result; The first feature fusion result is input into the initial language sub-model to obtain a first anomaly detection answer prediction result, and the first anomaly detection answer prediction result and the first anomaly detection answer label associated with the first sample image are used to update the initial text encoder and the initial language sub-model to obtain the universal text encoder and the universal language sub-model.

12. The system according to claim 7, characterized in that The first anomaly detection positioning label is a real positioning result of the device target in the first sample image; the first anomaly detection question text is a question text configured for requesting to obtain the anomaly detection result of the first sample image; The first anomaly detection answer label is a true target anomaly detection result of the first sample image, which is used to roughly determine the location of the anomaly.

13. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-modal large model anomaly detection method for power equipment described in any one of claims 1-6.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multi-modal large model anomaly detection method for electric power equipment according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • Anomaly detection method and device based on large visual language model

    CN117745680A