Fault monitoring method and device for power equipment and related equipment

Through the object detection model, the multimodal image of power equipment is detected and confidence analysis is solved, and the problem of inaccurate single-modal image monitoring is achieved, and more accurate fault monitoring is achieved.

CN120564136AActive Publication Date: 2025-08-29BEIJING TAIYUE TIANCHENG TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511054809.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-08-29
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

In the prior art, no matter which single-mode image is selected to analyze power equipment failures, it may lead to inaccurate conclusions, and single-mode images can only provide some useful information.

Method used

Through the object detection model, multiple types of single-modal images are detected, the prediction box is marked and the target prediction box is determined based on the confidence, and the useful information of different single-modal images is used for fault monitoring.

Benefits of technology

Improve the accuracy of power equipment fault monitoring, and provide more accurate fault monitoring results by integrating useful information from different single-mode images across images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564136A_ABST
    Figure CN120564136A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power equipment fault diagnosis, in particular to a power equipment fault monitoring method and device and related equipment. The method comprises the following steps: detecting each single-mode image in a plurality of different types of single-mode images about a power place through a target detection model to obtain a single-mode annotation image; according to the confidence coefficient, determining a target prediction frame in at least one prediction frame used for labeling the same target object in the multiple single-mode labeling images; determining a fault monitoring result of the corresponding power equipment according to the image content in the frame of the target prediction frame; according to the method and the device, the technical problem that an inaccurate conclusion can be obtained whether any single-mode image is selected to analyze whether the power equipment breaks down or not in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fault diagnosis of electric power equipment, and in particular to a fault monitoring method, apparatus and related equipment for electric power equipment. Background Art

[0002] In existing technical solutions for monitoring the fault status of power equipment, various information displayed by single-modal images of different modes of power equipment can be analyzed to determine whether the power equipment has failed; among them, single-modal images include: visible light images, infrared images, and fused images, etc.

[0003] According to practical experience, visible light images can provide rich detailed information, but they are not very effective at night or in low light conditions; infrared images can indicate the temperature of power equipment and perform well even at night and in bad weather conditions, but they lack sufficient detailed information; although fused images combine the advantages of the first two, the fusion process will inevitably lead to information loss, color distortion, and difficulty in balancing the contrast between visible light images and infrared images, resulting in the fused image possibly providing inaccurate information.

[0004] In summary, since the above-mentioned various single-modal images can usually only provide a part of the information useful for fault monitoring of power equipment in the application of fault monitoring of power equipment, no matter which single-modal image is selected to analyze whether a fault occurs in the power equipment, inaccurate conclusions may be obtained. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a fault monitoring method, device and related equipment for power equipment to solve the technical problem in the prior art that inaccurate conclusions may be obtained regardless of which single-modal image is selected to analyze whether a fault has occurred in the power equipment.

[0006] In a first aspect, the present application provides a method for fault monitoring of an electric power device, the method comprising: Performing detection on each of a plurality of unimodal images of different types related to the power site using a target detection model to obtain a unimodal annotated image; The unimodal annotated image is annotated with at least one prediction box and a confidence level corresponding to each prediction box; the image content within the prediction box includes a target object identified by the target detection model; and the confidence level indicates a predicted probability that the target object is an electric power device. determining, according to the confidence level, a target prediction frame from at least one prediction frame used to annotate the same target object in the plurality of unimodal annotated images; The corresponding fault monitoring result of the electric power equipment is determined according to the image content within the frame of the target prediction box.

[0007] In a second aspect, the present application provides a fault monitoring device for electric power equipment, the device comprising: a target detection module, a target screening module, and a fault diagnosis module; The target detection module is configured to detect each of a plurality of different types of single-modal images of the power site using a target detection model to obtain a single-modal annotated image; The unimodal annotated image is annotated with at least one prediction box and a confidence score corresponding to each prediction box; the image content within the prediction box includes the target object identified by the target detection model; and the confidence score indicates the predicted probability that the target object is an electric power device. The target screening module is configured to determine, according to the confidence level, a target prediction frame from at least one prediction frame used to annotate the same target object in the plurality of monomodal annotated images; The fault diagnosis module is used to determine the fault monitoring result of the corresponding power equipment based on the image content within the frame of the target prediction box.

[0008] In a third aspect, the present application provides an electronic device comprising a processor and a memory, wherein the memory is used to store an application program, and the processor runs or executes a software program stored in the memory so that the electronic device implements the above-mentioned method for fault monitoring of power equipment.

[0009] In a fourth aspect, the present application provides a computer-readable storage medium, which is used to store program codes executed by a processor, and the program codes are used to implement the above-mentioned fault monitoring method for power equipment.

[0010] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on an electronic device, the electronic device implements the above-mentioned fault monitoring method for power equipment.

[0011] Beneficial effects: The present application provides a method for fault monitoring of electric power equipment, the method comprising: detecting each of a plurality of different types of single-modal images of an electric power site using a target detection model to obtain a single-modal annotated image; wherein the single-modal annotated image is annotated with at least one prediction box and a confidence level corresponding to each prediction box; the image content within the box of the prediction box includes a target object identified by the target detection model; the confidence level indicates a predicted probability that the target object is electric power equipment; based on the confidence level, determining a target prediction box in at least one prediction box used to annotate the same target object in a plurality of single-modal annotated images; and determining a fault monitoring result of the corresponding electric power equipment based on the image content within the box of the target prediction box.

[0012] In summary, the present application uses the prediction frame used to label power equipment as a carrier of "information useful for fault monitoring of power equipment", and uses the confidence of the prediction frame as a measure of the "usefulness" of the "information useful for fault monitoring of power equipment", to determine the most useful target prediction frame for labeling the same target object in multiple single-modal labeled images, and subsequently perform fault diagnosis on the power equipment based on the image content within the frame of the target prediction frame; since the present application can break the ideological limitation that "single-modal images can usually only provide part of the information useful for fault monitoring of power equipment", and integrate and apply part of the useful information that different single-modal images can provide respectively, it can solve the technical problem in the prior art that inaccurate conclusions may be obtained regardless of which single-modal image is selected to analyze whether the power equipment has a fault. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. The following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0014] Figure 1 A flow chart of a method for monitoring faults in power equipment provided in this application; FIG2( a ) is an example of a visible light image to be labeled that meets the input image size of the target detection model provided by the present application; FIG2( b ) is an example diagram of a fused image to be annotated that meets the size of an input image of an object detection model provided by the present application; FIG2( c ) is an example diagram of an infrared amplified image to be labeled that meets the size of the input image of the target detection model provided by the present application; FIG2( d ) is an example of an infrared image to be labeled that meets the size of the input image of the target detection model provided by the present application; Figure 3 (a) is an example of a visible light annotated image provided by this application; Figure 3 (b) is an example of a fused and annotated image provided by this application; Figure 3 (c) is an example of an infrared amplified and annotated image provided by this application; Figure 4 (a) is an example diagram of the visible light marked area provided by this application; Figure 4 (b) is an example diagram of the fused annotation area provided by this application; FIG5( a ) is an example diagram of the first prediction frame provided by this application; FIG5( b ) is an example diagram of the second prediction frame provided by this application; FIG5( c ) is an example diagram of the third prediction frame provided by this application; FIG5( d ) is an example diagram of the fourth prediction frame provided by this application; FIG6 (a) is an example diagram of an alternative first prediction frame provided by this application; FIG6( b ) is an example diagram of a third prediction frame corresponding to an alternative first prediction frame provided by the present application; FIG6 (c) is an example diagram of the first target prediction frame and the first prediction frame to be selected provided by this application; FIG6( d ) is an example diagram of a third prediction box corresponding to the first target prediction box and a third prediction box corresponding to the first prediction box to be selected provided by the present application; FIG6 (e) is an example diagram of an alternative second prediction frame provided by this application; FIG6( f ) is an example diagram of a first prediction frame corresponding to an alternative second prediction frame provided by the present application; FIG6 (g) is an example diagram of the second target prediction frame and the second prediction frame to be selected provided by this application; FIG6(h) is an example diagram of the first prediction frame corresponding to the second target prediction frame provided by the present application; Figure 7 This is an example of an infrared annotated image provided for this application; Figure 8 This is a schematic diagram of the structure of the fault monitoring device for electric power equipment provided in this application. DETAILED DESCRIPTION

[0015] In existing image-based technologies for monitoring fault conditions in power equipment at power facilities, different single-modal images offer varying advantages and disadvantages. These images can provide different types of information. For example, visible light images provide rich details in bright sunlight but are less effective at night or in low-light conditions. Infrared images excel at indicating the temperature of power equipment and are well-suited for nighttime and inclement weather, but lack detailed information. Fusion images, while combining the advantages of both visible light and infrared images, are prone to information loss, color distortion, and contrast imbalance during the fusion process, potentially leading to inaccurate information.

[0016] For example, visible light images can provide detailed information about power equipment in a power plant under sufficient lighting, but cannot provide detailed information about power equipment in a shaded area. Infrared images can provide temperature information about power equipment with abnormal surface temperatures due to abnormal operation, but cannot provide additional information about power equipment with normal surface temperatures or with little temperature change. Fusion images often suffer from information loss, color distortion, and contrast balance issues during the image fusion process, resulting in the fusion image being unable to accurately provide detailed information about power equipment in a power plant under sufficient lighting, nor providing temperature information about power equipment with abnormal surface temperatures due to abnormal operation. In summary, since the above-mentioned single-modal images can usually only provide a portion of the information useful for power equipment fault monitoring applications, no matter which single-modal image is selected to analyze whether a power equipment fault has occurred, an inaccurate conclusion may be obtained.

[0017] In order to solve the above technical problems, the present application provides a technical solution for fault monitoring of power equipment based on multimodal images, which can break the limitation of the idea that "single-modal images can usually only provide a part of the information useful for fault monitoring of power equipment", so as to integrate and apply some of the useful information that different single-modal images can provide, and take a different approach to realize the "cross-image" application of "information useful for fault monitoring of power equipment". In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0018] First, the present application provides a method for fault monitoring of electric equipment, such as Figure 1 As shown, Figure 1This is a flow chart of the fault monitoring method for power equipment provided in this application. The method includes: S110 to S130, and the details are as follows: S110 : Detecting each of a plurality of different types of unimodal images of a power plant using an object detection model to obtain a unimodal labeled image.

[0019] Among them, the unimodal annotated image is annotated with at least one prediction box and a confidence level corresponding to each prediction box; the image content within the prediction box includes the target object identified by the target detection model; and the confidence level indicates the predicted probability that the target object is an electrical device.

[0020] Specifically, in an embodiment of the present application, the image contents displayed by multiple different types of single-modal images of the power site are the same or corresponding; for example, if the types of single-modal images include visible light images and infrared images, then the above-mentioned visible light images and infrared images should be images taken at the same shooting location with the same shooting angle and shooting distance. Since the visible light image and the infrared image are images taken at the same shooting location with the same shooting angle and shooting distance, the image content of the obtained visible light image and the image content of the infrared image are corresponding.

[0021] It should be noted that in practice, the size of a visible light image is typically larger than that of an infrared image. Therefore, the image content displayed by a visible light image is generally greater than that displayed by an infrared image. Therefore, if the types of unimodal images of a power plant are visible light and infrared, the image content displayed by these two unimodal images is "corresponding." In this embodiment of the present application, the target detection model is a neural network model used to identify and classify target objects included in the image, such as the YOLO (You Only Look Once) model. The target objects are determined based on the requirements for training the YOLO model. If the purpose of training the YOLO model is to identify bird types, the target objects are various birds. In this embodiment of the present application, the target detection model is used to identify and classify electrical equipment in a power plant. The types of electrical equipment that can be identified and classified are determined based on the training sample set used to train the target detection model. In practice, after multiple unimodal images to be identified and classified are sequentially input into the target detection model, the target detection model sequentially outputs recognition and classification results for each of the multiple unimodal images.

[0022] The recognition and classification results include: the location, confidence, type and type probability of the prediction box. The prediction box is usually a rectangular box, which is used to mark the target object in the monomodal image in the form of framing the target object. Therefore, the prediction box also has the function of positioning. In the embodiment of this application, the target object refers to the power equipment. In actual operation, the target detection model outputs the center coordinates of the prediction box. And the length and width of the prediction box , when the center coordinate of the prediction box And the length and width of the prediction box After determination, the prediction box can be drawn in the unimodal image. Type refers to the "type of electric equipment" to which the electric equipment belongs. Confidence indicates the predicted probability that the target object is an electric equipment; type probability indicates the predicted probability that the electric equipment belongs to a "type of electric equipment"; confidence and type probability are both numbers between 0 and 1; the closer the confidence is to 1, the more likely the target detection model believes that the prediction box has labeled the electric equipment clearly enough; the closer the type probability is to 1, the more likely the target detection model believes that the electric equipment belongs to the corresponding type of electric equipment. After the recognition and classification results of each unimodal image are determined according to the target detection model, the unimodal labeled image corresponding to each unimodal image can be determined based on the recognition and classification results.

[0023] In one implementation, the types of single-modality images include: visible light images, fused images, and infrared amplified images; before S110, the method further includes: steps (1) to (3), the details of which are as follows: Step (1): Take infrared images and visible light images of the power site to be annotated at the same shooting location, at the same shooting angle and shooting distance.

[0024] Specifically, in the embodiment of the present application, the visible light image to be annotated and the infrared image to be annotated are images taken by a visible light camera and an infrared camera at the same shooting location with the same shooting angle and shooting distance, respectively.

[0025] Step (2): Fusing the infrared image to be labeled and the visible light image to be labeled to obtain a fused image to be labeled.

[0026] Step (3): Based on the size of the infrared image to be annotated, the size of the visible light image to be annotated, the image content of the visible light image to be annotated, and the image content of the infrared image to be annotated, background filling is performed at the edge of the infrared image to be annotated to obtain the infrared amplified image to be annotated.

[0027] The size of the infrared amplified image to be annotated is the same as the size of the visible light image to be annotated; the relative position of the image content of the infrared image to be annotated in the infrared amplified image to be annotated is the same as the relative position of the image content of the infrared image to be annotated in the visible light image to be annotated.

[0028] Specifically, in actual operation, since the target detection model requires the size of the input image to be fixed when processing images, if the sizes of multiple input images to be labeled are inconsistent, the connection weights will not match, and the output results cannot be calculated. Therefore, before inputting multiple single-modal images such as infrared images to be labeled, visible light images to be labeled, and fused images to be labeled into the target detection model, it is necessary to adjust the multiple images to meet the size requirements of the target detection model for the input images; among them, the infrared amplified images to be labeled include the image content of the infrared images to be labeled.

[0029] Generally speaking, among the various types of images to be annotated, visible light images have the largest size. Therefore, to ensure that information from all images is not lost, the size of the visible light images to be annotated can be set to the size of the input image required by the target detection model. Before inputting the infrared images to be annotated into the target detection model, the size of the infrared images to be annotated is stretched to the same size as the visible light images. As shown in Figures 2(a), 2(b), 2(c), and 2(d), Figure 2(a) is an example of a visible light image to be annotated that meets the size of the input image for the target detection model provided by this application, Figure 2(b) is an example of a fused image to be annotated that meets the size of the input image for the target detection model provided by this application, Figure 2(c) is an example of an infrared amplified image to be annotated that meets the size of the input image for the target detection model provided by this application, and Figure 2(d) is an example of an infrared image to be annotated that meets the size of the input image for the target detection model provided by this application (i.e., a stretched infrared image to be annotated that meets the size of the input image for the target detection model provided by this application).

[0030] In the embodiment of the present application, the solid small squares and circular frames within the large square in Figure 2(a) represent target objects derived from the visible light image. The dashed small squares in Figures 2(b), 2(c), and 2(d) represent target objects derived from the infrared image, and the circular frame in Figure 2(b) represents target objects derived from the visible light image. It should be emphasized that in order to briefly illustrate the differences in image content between the visible light image to be labeled, the fused image to be labeled, the infrared amplified image to be labeled, and the infrared image to be labeled, the target objects are pre-displayed in the form of predicted frames in Figures 2(a), 2(b), 2(c), and 2(d). In actual operation, the visible light image to be labeled, the fused image to be labeled, the infrared amplified image to be labeled, and the infrared image to be labeled do not contain predicted frames. In actual operation, although stretching the infrared image to be labeled can make the stretched infrared image show more detailed information, it may also cause information loss or even image distortion. In order to retain the unstretched infrared image to be labeled, the present application constructs the infrared amplified image to be labeled, namely Figure 2(c). The infrared image shown in FIG2 (d) is already a stretched infrared image to be annotated. When the infrared image to be annotated shown in FIG2 (d) is not stretched, the size of the small dotted square is the same as the size of the small dotted square in FIG2 (b).

[0031] According to Figures 2(a), 2(c) and 2(d), the size of the infrared amplified image to be annotated shown in Figure 2(c) is the same as the size of the visible light image to be annotated shown in Figure 2(a), and the infrared amplified image to be annotated shown in Figure 2(c) includes the image content of the infrared image to be annotated when it is not stretched (i.e., the image content shown in the small dotted square).

[0032] S120: Determine, according to the confidence level, a target prediction box from at least one prediction box used to annotate the same target object in the plurality of unimodal annotated images.

[0033] Specifically, in order to solve the technical problem that "single-modal images can usually only provide part of the available information", which leads to the possibility of obtaining inaccurate conclusions when analyzing whether a fault occurs in power equipment regardless of any single-modal image, the embodiment of the present application extracts "information useful for fault monitoring of power equipment" from the "single-modal annotated image" corresponding to various "single-modal images", and then applies it comprehensively, that is, a "cross-image" comprehensive application; wherein, the image content within the prediction box in each single-modal annotated image is the "information useful for fault monitoring of power equipment" in the corresponding single-modal image, because the image content within the prediction box includes the power equipment to be monitored for faults. In actual operation, although the image content within the frame of each prediction box includes power equipment, that is, the image content within the frame belongs to "information useful for fault monitoring of power equipment", only the prediction box with the highest "usefulness" of the "information useful for fault monitoring of power equipment" that can be provided can obtain the most accurate fault monitoring result.

[0034] In an embodiment of the present application, the confidence of the prediction box is used as a measure of "usefulness", that is, for the power equipment located in the prediction box in different unimodal annotated images, the higher the confidence of the prediction box where the power equipment is located, the higher the "usefulness" of the "information useful for fault monitoring of the power equipment" that the prediction box can provide about the power equipment is; therefore, in an embodiment of the present application, the prediction box with the highest confidence among at least one prediction box used to annotate the same target object in multiple unimodal annotated images is determined as the target prediction box of the corresponding target object. In actual operation, the number of prediction boxes with "maximum confidence" can be one or more, that is, there may be multiple prediction boxes with equal and maximum confidence; for example, the number of prediction boxes currently used to mark the same target object is 4, and their corresponding confidences are 0.85, 0.74, 0.84 and 0.85, respectively, among which the confidence of 0.85 is the maximum confidence and there are two prediction boxes corresponding to the confidence of 0.85; in actual operation, when "there are multiple prediction boxes with equal and maximum confidence", a target prediction box can be selected from the above-mentioned multiple prediction boxes with the maximum confidence. The selection criteria can be determined according to actual needs, and this application does not make specific limitations on this.

[0035] The target prediction box is a prediction box that can provide "information useful for fault monitoring of power equipment" with the highest "usefulness" for fault monitoring of power equipment included in the image content within its frame.

[0036] In one implementation, S120 includes steps (4) to (6), the details of which are as follows: Step (4): Select a reference annotated image from multiple single-modal annotated images, and determine all prediction boxes in the reference annotated image as reference prediction boxes.

[0037] Specifically, according to actual conditions, the unimodal images used in the embodiments of the present application are all images of power sites, which usually include multiple power equipment. Therefore, each corresponding unimodal annotated image generally includes at least one prediction box, and multiple unimodal annotated images may include a large number of prediction boxes. Therefore, it is relatively difficult to select the prediction boxes that annotate the same target object from a large number of prediction boxes.

[0038] In actual operation, if the power equipment included in the image content within the frames of different prediction boxes in a single-modal annotated image is different types of power equipment, the prediction box used to label "the same power equipment" can be directly selected based on the "type" output by the target detection model. However, in actual applications, the power equipment included in the image content within the frames of different prediction boxes in a single-modal annotated image is usually the same type of power equipment. Therefore, it is impossible to accurately determine the prediction box used to label "the same power equipment" in any two single-modal annotated images based solely on the "type" output by the target detection model. In an embodiment of the present application, whether two prediction boxes are used to label "the same power equipment" is confirmed by determining the similarity between the image content within the frames of the two prediction boxes; if the similarity between the image content within the frames of the two prediction boxes is greater than or equal to a preset similarity threshold, it can be considered that the two prediction boxes are used to label "the same power equipment"; if the similarity between the image content within the frames of the two prediction boxes is less than the similarity threshold, it can be considered that the two prediction boxes are not used to label "the same power equipment". However, in actual operation, if we want to confirm the similarity of all prediction boxes included in all unimodal images pairwise, the workload will be very large.

[0039] In order to improve the efficiency of determining the prediction box used to label "the same power equipment", an embodiment of the present application determines one unimodal annotated image among multiple unimodal annotated images as a reference annotated image, and calculates the similarity between the image content within the frame of each prediction box in other unimodal annotated images except the basic annotated image and the image content within the frame of the prediction box in the unimodal image, so as to determine the prediction box used to label "the same power equipment" in multiple unimodal annotated images.

[0040] In an embodiment of the present application, the reference annotation image is determined based on actual needs. For example, if the fault monitoring at this time relies more on the surface details of the power equipment to determine whether a fault has occurred, the visible light annotation image can be determined as the reference annotation image. If the fault monitoring at this time relies more on the temperature information of the power equipment to determine whether a fault has occurred, the infrared annotation image can be determined as the reference annotation image.

[0041] When the reference annotation image is determined, the prediction box in the reference annotation image is the reference prediction box.

[0042] Step (5): among all prediction frames included in the remaining unimodal annotated images after removing the reference annotated image from the multiple unimodal annotated images, the prediction frames whose similarity between the image content within the frame of the prediction frame and the image content within the frame of the reference prediction frame is greater than or equal to the similarity threshold are determined as similar prediction frames.

[0043] Specifically, the similar prediction box is a prediction box that is determined in the embodiment of the present application through a similarity threshold in the remaining single-modal annotated images after removing the reference annotated image from multiple single-modal images, and is used to annotate "the same power equipment" together with the reference prediction box. In actual applications, the number of similar prediction boxes can be 0. When the number of similar prediction boxes is 0, it indicates that there is a target object in the reference annotated image that does not exist in other single-modal images; for example, if the reference annotated image is a visible light annotated image, since the image size of the visible light annotated image is generally larger than that of the infrared annotated image, the visible light annotated image will include more image content than the infrared annotated image, and therefore the visible light annotated image may include more prediction boxes than the infrared annotated image. Therefore, when the visible light annotated image is the reference annotated image, there may be no similar prediction box in the infrared annotated image that corresponds to a certain reference prediction box in the visible light annotated image.

[0044] Step (6): For each reference prediction frame, the prediction frame with the highest confidence between the reference prediction frame and the corresponding similar prediction frame is determined as the target prediction frame of the target object corresponding to the reference prediction frame.

[0045] Specifically, after determining similar prediction boxes and reference prediction boxes for annotating the same power equipment in multiple unimodal annotated images, the prediction box with the highest degree of usefulness in providing information useful for power equipment fault monitoring is selected based on its confidence level. This is the target prediction box for the power equipment. If a reference prediction box does not have a corresponding similar prediction box, the reference prediction box is determined as the target prediction box.

[0046] In one implementation, the types of single-modality annotated images include: visible light annotated images, fused annotated images, and infrared amplified annotated images; S120 includes: steps (7) to (14), the details of which are as follows: Step (7): Based on the image content of the infrared image to be annotated, the image areas in the visible light annotated image and the fused annotated image corresponding to the image content of the infrared image to be annotated are respectively determined as visible light annotated areas and fused annotated areas.

[0047] Specifically, as shown in Figures 3(a), 3(b) and 3(c), Figure 3(a) is an example image of a visible light annotated image provided in this application, Figure 3(b) is an example image of a fused annotated image provided in this application, and Figure 3(c) is an example image of an infrared amplified annotated image provided in this application. Since the image contents of different single-modal annotated images are different, the prediction boxes in the corresponding single-modal annotated images are not exactly the same.

[0048] In an embodiment of the present application, the focus of fault monitoring is placed on power equipment that can simultaneously determine detailed information and temperature information, that is, fault monitoring is only performed on multiple "same power equipment" that appear in both infrared images and visible light images. Therefore, the target prediction box that needs to be determined is also the target prediction box of the above-mentioned multiple "same power equipment".

[0049] In actual operation, in order to ensure that the power equipment corresponding to the determined target prediction frame appears simultaneously in the infrared annotated image and the visible light annotated image, it is necessary to intercept the visible light annotated image and the fused annotated image so that the image content of the intercepted visible light annotated area and the fused annotated area corresponds to the image content of the infrared annotated image, as shown in Figures 4 (a) and 4 (b). Figure 4 (a) is an example diagram of the visible light annotated area provided by this application, and Figure 4 (b) is an example diagram of the fused annotated area provided by this application. The image content within the rectangular frame located in the central area of ​​the image in Figures 4 (a) and 4 (b) is the visible light annotated area and the fused annotated area. It should be emphasized that in the process of fault monitoring of power equipment in power facilities, many images will be taken. For power equipment that has not been identified as a target prediction frame in this fault monitoring, the corresponding target prediction frame can be determined in fault monitoring based on other images, and then fault prediction can be performed based on this.

[0050] Step (8): Determine the prediction boxes included in the visible light annotation area as the first prediction box, determine the prediction boxes included in the fused annotation area as the second prediction box, and determine the prediction boxes included in the infrared amplified annotation image as the third prediction box.

[0051] Specifically, as shown in Table 1, Table 1 is a detailed list of prediction frames in each single-modal annotated image / annotated area provided by this application; wherein the first prediction frame The meaning is the visible light marked area included The first prediction box The first prediction box; the second prediction box The meaning is the fusion annotation area included The first The second prediction box; the third prediction box The meaning of the infrared amplified annotation image is The third prediction box The third prediction box, 、 、 、 、 and All are positive integers.

[0052] Table 1 Detailed list of prediction boxes in each single-modal annotated image / annotated area

[0053] It should be emphasized that the first prediction box and the second prediction box determined in the visible light annotated area and the fused annotated area in the embodiments of the present application are both "complete" prediction boxes, because the visible light annotated area and the fused annotated area are image regions captured from the visible light annotated image and the fused annotated image, and only include part of the image content of the visible light annotated image and the fused annotated image, and usually only include part of the prediction boxes of all the prediction boxes in the visible light annotated image and the fused annotated image. According to the aforementioned statement that "therefore, it is necessary to capture the visible light annotated image and the fused annotated image so that the image content of the captured visible light annotated area and the fused annotated area corresponds to the image content of the infrared annotated image", it can be seen that the capture of the visible light annotated area and the fused annotated area in the embodiments of the present application is based on "the image content of the captured annotated image is the same as the image content of the infrared annotated image". Since the position of the prediction box in the visible light annotated image and the fused annotated image may appear anywhere in the image in actual operation, in some cases, truncated prediction boxes may appear in the visible light annotated area and the fused annotated area captured from the visible light annotated image and the fused annotated image. In an embodiment of the present application, since fault monitoring is performed based on the similarity between prediction boxes in different annotated images and / or different annotated areas, the prediction box used to compare the similarity must have a complete "image content within the prediction box". If the "truncated prediction box" is used to calculate the similarity, the calculated similarity will inevitably be very low, which will seriously affect the determination of the fault monitoring results. Therefore, in the process of determining the first prediction box and the second prediction box, the selected prediction boxes are all "complete" prediction boxes.

[0054] Step (9): For each third prediction frame, the first prediction frame in which the similarity between the image content within the frame of all first prediction frames and the image content within the frame of the third prediction frame is greater than or equal to the similarity threshold is determined as an alternative first prediction frame.

[0055] Specifically, in an embodiment of the present application, as shown in Figures 5(a), 5(b), 5(c) and 5(d), Figure 5(a) is an example diagram of the first prediction box provided by the present application, Figure 5(b) is an example diagram of the second prediction box provided by the present application, Figure 5(c) is an example diagram of the third prediction box provided by the present application, and Figure 5(d) is an example diagram of the fourth prediction box provided by the present application. The image area in the large solid box in Figure 5 (a) represents the visible light annotated area, the small solid box in the large solid box in Figure 5 (a) represents the prediction box that comes from the visible light image and is located in the visible light annotated area, and the circular box in the large box in Figure 5 (a) represents the prediction box that comes from the visible light image and is located outside the visible light annotated area; the image area in the large dotted box in Figure 5 (b) represents the fused annotated area, the small dotted box in Figure 5 (b) represents the prediction box of the image part that fuses the visible light image information and the infrared image information (in order to intuitively represent the fusion, the small dotted box of the infrared image prediction box is still used), that is, the prediction box located in the fused annotated area, and the circular box in Figure 5 (b) represents the prediction box that comes from the visible light image and is located outside the fused annotated area. Among them, the prediction boxes in Figures 5 (a), 5 (b), 5 (c) and 5 (d) are filled in. 、 、 、 Correspondingly, it represents the confidence level corresponding to the predicted box.

[0056] In actual operation, due to the infrared amplification annotation image The image content includes the unstretched infrared image, so in the process of determining the target prediction frame, the visible light annotation area is first The first predicted box and infrared amplified annotation image in The first round of voting is performed on the third prediction box to filter out the prediction box used to label "the same power equipment".

[0057] ; Where, Indicates the visible light annotation area Included The first The first prediction box; Indicates infrared amplified annotation image Included The third prediction box A third prediction box; Indicates infrared amplified annotation image The third prediction box included The main voting is to use infrared amplification to mark the image. Included The third prediction frame is used as the benchmark; Indicates the area marked for visible light and infrared amplified annotation images Voting performed on two unimodal annotated images / annotated regions.

[0058] In actual operation, infrared amplified annotation images Included For each third prediction frame in the third prediction frames, determine the first prediction frame in which the similarity between the image content within the frame of all first prediction frames and the image content within the frame of the third prediction frame is greater than or equal to the similarity threshold. According to the above discussion, if "the similarity between the image content within the frame of the two prediction frames is greater than or equal to the similarity threshold", the two prediction frames are considered to be used to mark "the same power equipment"; as shown in Figures 6 (a) and 6 (b), Figure 6 (a) is an example diagram of the alternative first prediction frame provided by this application, and Figure 6 (b) is an example diagram of the third prediction frame corresponding to the alternative first prediction frame provided by this application. Figure 6 (a) is marked with the alternative first prediction frame obtained after step (9), and Figure 6 (b) is marked with the third prediction frame corresponding to the alternative first prediction frame obtained after step (9).

[0059] In the embodiment of the present application, the visible light marking area Infrared amplified and annotated images The third prediction box in is used to mark the first prediction box of "the same power equipment" and is determined as the candidate first prediction box. It should be emphasized that the number of candidate first prediction boxes may be multiple, because the image content of the visible light area image and the image content of the infrared amplified annotation image may include multiple "the same power equipment", so the corresponding visible light annotation area and infrared amplified annotation images The image may include multiple groups of candidate first prediction boxes and third prediction boxes for marking "the same electric power device", and each group of candidate first prediction boxes and third prediction boxes corresponds to one "the same electric power device".

[0060] Step (10): For each third prediction frame, the prediction frame with the highest confidence among the third prediction frame and the corresponding candidate first prediction frame is determined as the first target prediction frame.

[0061] Specifically, after determining multiple groups of alternative first prediction boxes and third prediction boxes for marking "the same power equipment", the prediction box with the highest "usefulness" of "information useful for fault monitoring of power equipment" that can be provided can be selected from each group of alternative first prediction boxes and third prediction boxes according to the confidence level, that is, the first target prediction box.

[0062] In actual operation, infrared amplified annotation images The remaining third prediction boxes that have not been determined as the first target prediction boxes are determined as the third prediction boxes to be selected for subsequent voting.

[0063] As shown in Table 2, Table 2 is a detailed list of prediction boxes in each single-modal annotated image / annotated area after the first round of voting provided by this application, which mainly lists the details of the alternative first prediction box, the first prediction box to be selected, and other prediction boxes.

[0064] Table 2 Detailed list of prediction boxes in each single-modal annotated image / annotated area

[0065] As shown in Figures 6(c) and 6(d), Figure 6(c) is an example diagram of the first target prediction frame and the first prediction frame to be selected provided by the present application, and Figure 6(d) is an example diagram of the third prediction frame corresponding to the first target prediction frame and the third prediction frame corresponding to the first prediction frame to be selected provided by the present application; Figure 6(c) is marked with the first target prediction frame obtained after step (10), the alternative first prediction frame remaining after removing the first target prediction frame, and the first prediction frame to be selected, and the first prediction frame to be selected is the prediction frame remaining after removing the alternative first prediction frame and the first target prediction frame in the visible light annotation area; Figure 6(d) is marked with the third prediction frame corresponding to the alternative first prediction frame obtained after step (10), the third prediction frame corresponding to the remaining alternative first prediction frame, and the third prediction frame corresponding to the first prediction frame to be selected.

[0066] It should be emphasized that although in the example shown in Figure 6 (c), the "first target prediction box" is derived from the first prediction box, in the embodiment of the present application, since the determined first target prediction box is determined in the "third prediction box" and the "alternative first prediction box corresponding to the third prediction box", there is a situation where the "third prediction box" is determined as the first target prediction box; if there is a situation where the "third prediction box" is determined as the first target prediction box, it means that during the fault monitoring process, the image content within the border of the third prediction box in the infrared amplified annotation image is clearer. This is also the goal of the technical solution of the embodiment of the present application, that is, under the premise of acknowledging the existence of various operational errors such as image acquisition errors, the prediction box that can most clearly represent the power equipment is determined, so as to determine the most accurate fault monitoring result based on the prediction box.

[0067] Step (11): Determine as candidate second prediction frames any second prediction frame whose image content within its frame is greater than or equal to a similarity threshold to the image content within the frame of the candidate first prediction frame. The candidate first prediction frame is the prediction frame remaining after excluding the candidate first prediction frame and the first target prediction frame (if the first target prediction frame exists in the visible light annotated area) from the visible light annotated area.

[0068] Specifically, in the embodiment of the present application, after the first round of voting, the visible light marking area is selected. and fused annotation areas The first prediction box and the second prediction box to be selected are voted for in the second round to select the prediction box for labeling “the same power equipment”.

[0069] ; Where, Indicates the visible light annotation area The A first prediction box to be selected; Indicates the fused annotation area Included The first In actual operation, for the visible light annotation area For each first prediction frame to be selected, determine all second prediction frames whose image contents within the frames are greater than or equal to a similarity threshold value with respect to the image contents within the frame of the first prediction frame to be selected.

[0070] In the embodiment of the present application, the marked area is fused Medium and visible light marking area The first prediction frame to be selected is used to mark the second prediction frame of "the same power equipment" and is determined to be the candidate second prediction frame. As shown in Figures 6 (e) and 6 (f), Figure 6 (e) is an example diagram of the candidate second prediction frame provided by this application, and Figure 6 (f) is an example diagram of the first prediction frame corresponding to the candidate second prediction frame provided by this application; Figure 6 (e) is marked with the candidate second prediction frame obtained after step (11), and Figure 6 (f) is marked with the first prediction frame corresponding to the candidate second prediction frame obtained after step (11), the candidate first prediction frame, the first target prediction frame, and the remaining candidate first prediction frames.

[0071] Step (12): For each candidate first prediction frame, the prediction frame with the highest confidence between the candidate first prediction frame and the candidate second prediction frame is determined as the second target prediction frame.

[0072] Specifically, after determining multiple groups of candidate first prediction boxes and alternative second prediction boxes for marking "the same power equipment", the prediction box with the highest "usefulness" of "information useful for fault monitoring of power equipment" that can be provided can be selected from each group of candidate first prediction boxes and alternative second prediction boxes according to the confidence level, that is, the second target prediction box.

[0073] As shown in Table 3, Table 3 is a detailed list of prediction boxes in each single-modal annotated image / annotated area after the second round of voting provided by this application, which mainly states the details of the alternative first prediction box, alternative first prediction box, and other prediction boxes.

[0074] Table 3 Detailed list of prediction boxes in each single-modal annotated image / annotated area

[0075] As shown in Figures 6(g) and 6(h), Figure 6(g) is an example diagram of the second target prediction box and the second prediction box to be selected provided by the present application. The second prediction box to be selected is the second prediction box remaining after removing the alternative second prediction box and the second target prediction box from the fusion annotation area; Figure 6(h) is an example diagram of the first prediction box corresponding to the second target prediction box provided by the present application; Figure 6(g) is marked with the second target prediction box and the second prediction box to be selected obtained after step (12), and Figure 6(h) is marked with the first prediction box corresponding to the alternative second prediction box, the first prediction box corresponding to the second target prediction box, the remaining first prediction box to be selected, the alternative first prediction box and the first target prediction box obtained after step (12).

[0076] It should be emphasized that although in the example shown in Figure 6 (g) the "second target prediction box" comes from the alternative second prediction box, in the embodiment of the present application, since the determined second target prediction box is determined in the "first prediction box to be selected" and the "alternative second prediction box", there is a situation where the "alternative second prediction box" is determined as the second target prediction box.

[0077] Step (13): For each candidate third prediction frame, the second prediction frame whose similarity between the image content within the frame of all candidate second prediction frames and the image content within the frame of each candidate third prediction frame is less than the similarity threshold is determined as the third target prediction frame.

[0078] Among them, the second prediction box to be selected is the second prediction box remaining after removing the alternative second prediction box and the second target prediction box (if there is a second target prediction box in the fused annotation area) in the fused annotation area; the third prediction box to be selected is the remaining third prediction box that has not been determined as the first target prediction box.

[0079] Specifically, in the embodiment of the present application, after the second round of voting, the fusion marked area and infrared amplified annotation images The second prediction box to be selected and the third prediction box to be selected are voted for in the third round to select the prediction box for labeling “the same power equipment”.

[0080] ; Where, Indicates the fused annotation area The A second prediction box to be selected; Indicates infrared amplified annotation image The In actual operation, for infrared amplified annotation images For each of the third prediction frames to be selected, determine the second prediction frames to be selected whose similarities between the image contents within the frames of all the second prediction frames to be selected and the image contents within the frames of each of the third prediction frames to be selected are less than a similarity threshold.

[0081] In the embodiment of the present application, the fusion marked area Infrared amplified and annotated images The third candidate prediction box in the image is used to mark the second candidate prediction box for "the same power equipment" and is determined as the third target prediction box. In this embodiment of the present application, the image information in the fused annotation area of ​​the fused annotation image and the image information in the infrared amplified annotation image both contain image information derived from the infrared image. Therefore, the similarity between the second candidate prediction box and the third candidate prediction box corresponding to the same target object in the fused annotation area and the infrared amplified annotation image must be very high (essentially 100%). Therefore, when determining the third target prediction box, the screening principle is not "similarity greater than or equal to the similarity threshold."

[0082] It should be emphasized that the reason why the above-mentioned "the similarity between the second prediction frame to be selected and the third prediction frame to be selected must be very high" is that the similarity between the second prediction frame to be selected and the third prediction frame to be selected "corresponding to the same target object" is very high. However, according to Figure 6 (d) and Figure 6 (g), the remaining second prediction frames to be selected and the third prediction frames to be selected do not all correspond to the same target object, so the similarity between the remaining second prediction frames to be selected and the third prediction frames to be selected is not all very high. Only the similarity between the second prediction frames to be selected and the third prediction frames to be selected "corresponding to the same target object" is very high. Therefore, the role of step (13) is to "collect" the remaining second prediction frames and the third prediction frames to be selected in the above-mentioned comparison process to jointly vote for the next time. Therefore, the screening principle for determining the third target prediction frame is "less than the similarity threshold" rather than "greater than the similarity threshold", because the second prediction frame or the third prediction frame "greater than the similarity threshold" has been correspondingly determined as the second target prediction frame and the third target prediction frame.

[0083] Step (14): Determine the first target prediction box, the second target prediction box, and the third target prediction box as the target prediction box.

[0084] Specifically, after three rounds of voting, the visible light marking area can be marked , fusion annotation area and infrared amplified annotation images The first target prediction frame, the second target prediction frame and the third target prediction frame used to mark “the same power equipment” are determined separately, and then they are integrated to obtain the target prediction frame. , fusion annotation area and infrared amplified annotation images The prediction box in the image can mark out the power equipment in the image content of the corresponding visible light image, fused image and infrared amplified image, so usually only the visible light annotation area is needed. , fusion annotation area and infrared amplified annotation images The accurate target detection frame can be determined.

[0085] In one implementation, the type of the single-modal annotated image further includes: an infrared annotated image; after step (14), the method further includes: steps (15) to (16), the details of which are as follows: Step (15): Determine the prediction box included in the infrared annotated image as the fourth prediction box.

[0086] Specifically, if Figure 7 As shown, Figure 7This is an example image of the infrared annotated image provided in this application, and the prediction box on it is the fourth prediction box.

[0087] Step (16): All fourth prediction frames whose image contents within the borders have a similarity less than a similarity threshold with the image contents within the borders of the target prediction frame are added to the target prediction frame.

[0088] Specifically, in the embodiment of the present application, the fourth prediction frame of the infrared annotated image mainly plays the role of "checking for missing and filling in the gaps", because the infrared annotated image in the embodiment of the present application is a single-modal annotated image corresponding to the stretched infrared image, which may include annotated visible light annotated area. , fusion annotation area and infrared amplified annotation images Therefore, the fourth prediction frame remaining after removing the prediction frame used to annotate "the same power equipment" with the target prediction frame in the infrared annotated image is the prediction frame accidentally obtained due to the stretching of the infrared image.

[0089] S130: Determine the fault monitoring result of the corresponding power equipment according to the image content within the frame of the target prediction box.

[0090] Specifically, after the target prediction frame is determined, the fault monitoring result of the corresponding power equipment can be determined based on the image content within the frame of the target prediction frame; among them, the technical solution of determining whether the corresponding power equipment has a fault based on the image content within the frame can be determined based on actual conditions, and this application does not make specific limitations on this.

[0091] In summary, the embodiment of the present application uses the prediction frame used to label the power equipment as a carrier of "information useful for fault monitoring of the power equipment", and uses the confidence of the prediction frame as a measure of the "usefulness" of the "information useful for fault monitoring of the power equipment", to determine the most useful target prediction frame for labeling the same target object in multiple single-modal labeled images, and subsequently perform fault diagnosis on the power equipment based on the image content within the frame of the target prediction frame.

[0092] Second, the present application provides a fault monitoring device for electric power equipment, such as Figure 8 As shown, Figure 8This is a schematic structural diagram of a fault monitoring device for electric power equipment provided in the present application, and the device includes: a target detection module 310, a target screening module 320, and a fault diagnosis module 330. The target detection module 310 is used to detect each of a plurality of different types of single-modal images of electric power sites through a target detection model to obtain a single-modal annotated image. The single-modal annotated image is annotated with at least one prediction box and a confidence level corresponding to each prediction box; the image content within the box of the prediction box includes the target object identified by the target detection model; the confidence level indicates the predicted probability that the target object is an electric power device; the target screening module 320 is used to determine the target prediction box in at least one prediction box used to annotate the same target object in a plurality of single-modal annotated images according to the confidence level; the fault diagnosis module 330 is used to determine the corresponding fault monitoring result of the electric power device according to the image content within the box of the target prediction box.

[0093] In one implementation, the target screening module 320 is further used to select a reference annotated image from a plurality of unimodal annotated images, and determine all prediction boxes in the reference annotated image as reference prediction boxes; the target screening module 320 is further used to determine, as similar prediction boxes, prediction boxes whose similarity between the image content within the borders of the prediction boxes and the image content within the borders of the reference prediction boxes is greater than or equal to a similarity threshold among all prediction boxes included in the unimodal annotated images remaining after removing the reference annotated image from the plurality of unimodal annotated images; the target screening module 320 is further used to, for each reference prediction box, determine the prediction box with the highest confidence between the reference prediction box and the corresponding similar prediction box as the target prediction box of the target object corresponding to the reference prediction box.

[0094] In one implementation, the types of single-modal images include visible light images, fused images, and infrared-amplified images. The apparatus further includes a single-modal image module; the single-modal image module is configured to capture an infrared image and a visible light image of a power facility at the same location, angle, and distance; the single-modal image module is further configured to fuse the infrared image and the visible light image to obtain a fused image; and the single-modal image module is further configured to perform background filling at the edges of the infrared image to obtain an infrared-amplified image based on the size of the infrared image to be labeled, the size of the visible light image to be labeled, and the image content of the visible light image to be labeled. The size of the infrared-amplified image to be labeled is the same as the size of the visible light image to be labeled; and the relative position of the image content of the infrared image to be labeled in the infrared-amplified image to be labeled is the same as the relative position of the image content of the infrared image to be labeled in the visible light image to be labeled.

[0095] In one implementation, the types of single-modal annotated images include: visible light annotated images, fused annotated images, and infrared-amplified annotated images. The target screening module 320 is further configured to, based on the image content of the infrared image to be annotated, determine image regions in the visible light annotated image and the fused annotated image corresponding to the image content of the infrared image to be annotated as visible light annotated regions and fused annotated regions, respectively. The target screening module 320 is further configured to determine all prediction boxes included in the visible light annotated regions as first prediction boxes, all prediction boxes included in the fused annotated regions as second prediction boxes, and all prediction boxes included in the infrared-amplified annotated images as third prediction boxes. The target screening module 320 is further configured to, for each third prediction box, determine as a candidate first prediction box a first prediction box whose image content within the frame of all first prediction boxes has a similarity greater than or equal to a similarity threshold with the image content within the frame of the third prediction box. The target screening module 320 is further configured to, for each third prediction box, determine as a candidate first prediction box a prediction box with the highest confidence level among the third prediction box and the corresponding candidate first prediction box as a first target prediction box.

[0096] In one implementation, the target screening module 320 is further configured to determine, as candidate second prediction frames, second prediction frames whose image contents within the frames are similar to or equal to the similarity threshold value with the image contents within the frames of the candidate first prediction frames; wherein the candidate first prediction frames are the prediction frames remaining after removing the candidate first prediction frames and the first target prediction frames from the visible light annotation area; the target screening module 320 is further configured to, for each candidate first prediction frame, determine the prediction frame with the largest confidence between the candidate first prediction frame and the candidate second prediction frame as the second target prediction frame; the target screening module 320 , and is further used to determine, for each candidate third prediction frame, the second prediction frame in all candidate second prediction frames, whose image content within the frame has a similarity less than a similarity threshold with the image content within the frame of each candidate third prediction frame, as the third target prediction frame; wherein the candidate second prediction frame is the second prediction frame remaining after removing the candidate second prediction frame and the second target prediction frame from the fused annotation area; the candidate third prediction frame is the remaining third prediction frame that has not been determined as the first target prediction frame; the target screening module 320 is further used to determine the first target prediction frame, the second target prediction frame, and the third target prediction frame as the target prediction frame.

[0097] In one implementation, the type of single-modal annotated image also includes: an infrared annotated image; the target screening module 320 is further used to determine the prediction box included in the infrared annotated image as a fourth prediction box; the target screening module 320 is further used to add the fourth prediction boxes in all fourth prediction boxes whose similarity between the image content within the frame and the image content within the frame of the target prediction box is less than a similarity threshold to the target prediction box.

[0098] Third, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, steps S110 to S130 provided in the above embodiment are implemented.

[0099] Fourth, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, steps S110 to S130 of the above embodiment are executed.

[0100] Fifth, the computer program product provided in this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. For specific implementation, please refer to steps S110 to S130 of the method embodiment, which will not be repeated here.

[0101] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0102] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

Claims

1. A method for monitoring faults of electric power equipment, characterized in that: The method comprises: Performing detection on each of a plurality of unimodal images of different types related to the power site using a target detection model to obtain a unimodal annotated image; The unimodal annotated image is annotated with at least one prediction box and a confidence score corresponding to each prediction box; the image content within the prediction box includes the target object identified by the target detection model; and the confidence score indicates the predicted probability that the target object is an electric power device. determining, according to the confidence level, a target prediction frame from at least one prediction frame used to annotate the same target object in the plurality of unimodal annotated images; The fault monitoring result of the corresponding power equipment is determined according to the image content within the frame of the target prediction box.

2. The method according to claim 1, characterized in that The step of determining, based on the confidence level, a target prediction frame from at least one prediction frame used to annotate the same target object in the plurality of unimodal annotated images comprises: Selecting a reference annotated image from the plurality of single-modality annotated images, and determining all prediction boxes in the reference annotated image as reference prediction boxes; Determining, among all prediction frames included in the remaining unimodal annotated images after removing the reference annotated image from the plurality of unimodal annotated images, prediction frames whose image content within the frame is similar to the image content within the frame of the reference prediction frame is greater than or equal to a similarity threshold as similar prediction frames; For each reference prediction frame, the prediction frame with the largest confidence among the reference prediction frame and the corresponding similar prediction frame is determined as the target prediction frame of the target object corresponding to the reference prediction frame.

3. The method according to claim 2, characterized in that The single-modal image includes: a visible light image, a fused image, and an infrared-amplified image; before determining, based on the confidence level, the target prediction frame in at least one prediction frame used to annotate the same target object in the plurality of single-modal annotated images, the method further includes: At the same shooting location, at the same shooting angle and shooting distance, shoot an infrared image and a visible light image to be labeled of the power site; Performing a fusion process on the infrared image to be labeled and the visible light image to be labeled to obtain a fused image to be labeled; Performing background filling at the edge of the infrared image to be labeled according to the size of the infrared image to be labeled, the size of the visible light image to be labeled, the image content of the visible light image to be labeled, and the image content of the infrared image to be labeled to obtain an infrared amplified image to be labeled; Among them, the size of the infrared amplified image to be labeled is the same as the size of the visible light image to be labeled; the relative position of the image content of the infrared image to be labeled in the infrared amplified image to be labeled is the same as the relative position of the image content of the infrared image to be labeled in the visible light image to be labeled.

4. The method according to claim 3, characterized in that The single-modal annotated image includes: a visible light annotated image, a fused annotated image, and an infrared amplified annotated image; for each reference prediction frame, determining the prediction frame with the highest confidence between the reference prediction frame and the corresponding similar prediction frame as the target prediction frame of the target object corresponding to the reference prediction frame includes: According to the image content of the infrared image to be annotated, image regions corresponding to the image content of the infrared image to be annotated in the visible light annotated image and in the fused annotated image are respectively determined as visible light annotated regions and fused annotated regions; Determining the prediction boxes included in the visible light annotation area as the first prediction box, determining the prediction boxes included in the fused annotation area as the second prediction box, and determining the prediction boxes included in the infrared amplified annotation image as the third prediction box; For each third prediction frame, determining, among all first prediction frames, those first prediction frames whose image contents within the frames are greater than or equal to a similarity threshold value to the image contents within the frames of the third prediction frame as candidate first prediction frames; For each third prediction frame, the prediction frame with the highest confidence among the third prediction frame and the corresponding candidate first prediction frame is determined as the first target prediction frame.

5. The method according to claim 4, characterized in that The step of determining, for each reference prediction frame, the prediction frame with the highest confidence between the reference prediction frame and the corresponding similar prediction frames as the target prediction frame of the target object corresponding to the reference prediction frame further includes: Determine, as candidate second prediction frames, those second prediction frames in which the similarity between the image content within the frame and the image content within the frame of the first prediction frame to be selected is greater than or equal to the similarity threshold; The to-be-selected first prediction frame is the prediction frame remaining after removing the candidate first prediction frame and the first target prediction frame from the visible light annotation area; For each candidate first prediction frame, determining the prediction frame with the highest confidence between the candidate first prediction frame and the candidate second prediction frame as the second target prediction frame; For each third prediction frame to be selected, determining as the third target prediction frame the second prediction frame whose similarity between the image content within the frame of all second prediction frames to be selected and the image content within the frame of each third prediction frame to be selected is less than the similarity threshold; The second prediction box to be selected is the second prediction box remaining after removing the candidate second prediction box and the second target prediction box from the fused annotation area; the third prediction box to be selected is the remaining third prediction box that has not been determined as the first target prediction box; The first target prediction box, the second target prediction box and the third target prediction box are determined as the target prediction box.

6. The method according to claim 5, characterized in that The single-modal annotated image further includes an infrared annotated image; after determining the first target prediction box, the second target prediction box, and the third target prediction box as the target prediction box, the method further includes: Determining the prediction frame included in the infrared annotated image as a fourth prediction frame; All the fourth prediction frames whose image contents within the frames are less than the similarity threshold value with the image contents within the frame of the target prediction frame are added to the target prediction frame.

7. A fault monitoring device for electric power equipment, characterized in that: The device comprises: a target detection module, a target screening module and a fault diagnosis module; The target detection module is configured to detect each of a plurality of different types of single-modal images of the power site using a target detection model to obtain a single-modal annotated image; The unimodal annotated image is annotated with at least one prediction box and a confidence score corresponding to each prediction box; the image content within the prediction box includes the target object identified by the target detection model; and the confidence score indicates the predicted probability that the target object is an electric power device. The target screening module is configured to determine, according to the confidence level, a target prediction frame from at least one prediction frame used to annotate the same target object in the plurality of monomodal annotated images; The fault diagnosis module is used to determine the fault monitoring result of the corresponding power equipment based on the image content within the frame of the target prediction box.

8. An electronic device, characterized in that: The electronic device includes a processor and a memory, the memory is used to store an application program, and the processor runs or executes a software program stored in the memory so that the electronic device implements the fault monitoring method for electric power equipment according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program codes executed by a processor, and the program codes are used to implement the fault monitoring method for electric power equipment according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device implements the fault monitoring method for electric power equipment according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Infrared image fault detection method and device for power equipment

    CN111986172A

  • System equipment state monitoring method and related device

    CN117994618A

  • Equipment fault detection method, robot and electronic equipment

    CN119354262A

  • Multi-wavelength video image fire detecting system

    US20090315722A1