A fault monitoring method and device of a power equipment and related equipment
By using a target detection model to detect and analyze the confidence of multiple monomodal images, and by comprehensively utilizing the information from multiple monomodal images of power equipment, the problem of inaccurate monomodal image monitoring is solved, and more accurate fault monitoring is achieved.
Patent Information
- Application Number
- CN202511054809.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-30
AI Technical Summary
In existing technologies, regardless of which single-modal image is selected for power equipment fault analysis, inaccurate conclusions may be obtained, as single-modal images can only provide partially useful information.
The target detection model is used to detect multiple types of monomodal images, the predicted bounding boxes are labeled and the confidence level is determined, and the useful information from different monomodal images is used for fault monitoring.
It improves the accuracy of power equipment fault monitoring by comprehensively utilizing useful information from multiple single-modal images, overcoming the limitations of single-modal images, and providing more accurate fault monitoring results.
Smart Images

Figure CN120564136B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power equipment fault diagnosis, and in particular to a power equipment fault monitoring method and device and related equipment. BACKGROUND
[0002] In the prior art, the various information displayed by the single-modal images of different modalities of the power equipment can be analyzed to determine whether the power equipment has failed. The single-modal images include visible light images, infrared images, and fusion images.
[0003] According to actual experience, visible light images can provide rich detailed information, but are not effective at night or in low light conditions. Infrared images can indicate the temperature of the power equipment and still perform well at night and in bad weather conditions, but lack sufficient detailed information. Although fusion images combine the advantages of the first two, information loss, color distortion, and difficulty in balancing the contrast between visible light images and infrared images inevitably occur during the fusion process, resulting in fusion images that may provide inaccurate information.
[0004] As can be seen from the above, since the various single-modal images described above can only provide part of the information useful for power equipment fault monitoring in the application of power equipment fault monitoring, it is possible to obtain inaccurate conclusions regardless of which single-modal image is selected to analyze whether the power equipment has failed. SUMMARY
[0005] Therefore, the purpose of the present application is to provide a power equipment fault monitoring method, device and related equipment to solve the technical problem in the prior art that inaccurate conclusions may be obtained regardless of which single-modal image is selected to analyze whether the power equipment has failed.
[0006] In a first aspect, the present application provides a power equipment fault monitoring method, which comprises:
[0007] detecting each of the single-modal images of different types in the plurality of single-modal images of the power site by a target detection model to obtain single-modal labeled images;
[0008] The single-modal labeled images are labeled with at least one prediction box and a confidence corresponding to each prediction box. The image content within the bounding box of the prediction box includes a target object recognized by the target detection model. The confidence indicates the predicted probability that the target object is a power equipment.
[0009] According to the confidence, at least one prediction box for labeling the same target object in a plurality of the single-modal labeled images is determined as a target prediction box.
[0010] According to the image content in the edge frame of the target prediction box, a fault monitoring result of the corresponding power equipment is determined.
[0011] In a second aspect, the present application provides a fault monitoring device for power equipment, the device comprising: a target detection module, a target screening module and a fault diagnosis module.
[0012] The target detection module is configured to detect each single-modal image in a plurality of single-modal images of different types about a power site by a target detection model to obtain a single-modal labeled image.
[0013] The single-modal labeled image is labeled with at least one prediction box and a confidence corresponding to each prediction box; the image content in the edge frame of the prediction box includes a target object recognized by the target detection model; and the confidence indicates a predicted probability that the target object is a power equipment.
[0014] The target screening module is configured to determine a target prediction box from at least one prediction box for labeling the same target object in a plurality of the single-modal labeled images according to the confidence.
[0015] The fault diagnosis module is configured to determine a fault monitoring result of the corresponding power equipment according to the image content in the edge frame of the target prediction box.
[0016] In a third aspect, the present application provides an electronic device, comprising a processor and a memory, the memory being configured to store an application program, and the processor being configured to run or execute a software program stored in the memory, so that the electronic device implements the above-mentioned fault monitoring method for power equipment.
[0017] In a fourth aspect, the present application provides a computer readable storage medium, which is configured to store program codes executed by a processor, and the program codes are configured to implement the above-mentioned fault monitoring method for power equipment.
[0018] In a fifth aspect, the present application provides a computer program product, which comprises computer instructions, and when the computer instructions are run on an electronic device, the electronic device implements the above-mentioned fault monitoring method for power equipment.
[0019] Advantages:
[0020] The application provides a fault monitoring method of a power equipment, and the method comprises the following steps: detecting each single-mode image in multiple single-mode images of different types about a power site by a target detection model to obtain a single-mode labeled image; wherein, at least one prediction box and a confidence corresponding to each prediction box are labeled on the single-mode labeled image; the image content in the edge frame of the prediction box comprises a target object recognized by the target detection model; the confidence indicates a prediction probability that the target object is the power equipment; according to the confidence, a target prediction box is determined from at least one prediction box for labeling the same target object in the multiple single-mode labeled images; and a fault monitoring result of the corresponding power equipment is determined according to the image content in the edge frame of the target prediction box.
[0021] As can be seen, the application takes the prediction box for labeling the power equipment as a carrier of the information useful for the fault monitoring of the power equipment, and takes the confidence of the prediction box as a measurement standard of the usefulness of the information useful for the fault monitoring of the power equipment, for determining the most useful target prediction box for labeling the same target object from the multiple single-mode labeled images, and subsequently performing fault diagnosis on the power equipment according to the image content in the edge frame of the target prediction box; since the application can break the thought limitation that the single-mode image can usually only provide part of the information useful for the fault monitoring of the power equipment, and integrate and apply the part of useful information respectively provided by different single-mode images, the technical problem that no matter which single-mode image is selected to analyze whether the power equipment has a fault, an inaccurate conclusion may be obtained in the prior art can be solved. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments of the application. The following drawings only show some embodiments of the application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0023] Figure 1 The flowchart of the fault monitoring method of the power equipment provided by the application is shown in the figure;
[0024] Fig. 2(a) is an example diagram of a visible light image to be labeled provided by the application, which conforms to the size of the input image of the target detection model;
[0025] Fig. 2(b) is an example diagram of a fusion image to be labeled provided by the application, which conforms to the size of the input image of the target detection model;
[0026] Fig. 2(c) is an example diagram of an infrared amplification image to be labeled provided by the application, which conforms to the size of the input image of the target detection model;
[0027] FIG. 2(d) is an example of an infrared image to be labeled provided by the present application, which has a size of an input image conforming to a target detection model;
[0028] FIG. 3(a) is an example of a visible light labeled image provided by the present application;
[0029] FIG. 3(b) is an example of a fusion labeled image provided by the present application;
[0030] FIG. 3(c) is an example of an infrared augmented labeled image provided by the present application;
[0031] FIG. 4(a) is an example of a visible light labeled region provided by the present application;
[0032] FIG. 4(b) is an example of a fusion labeled region provided by the present application;
[0033] FIG. 5(a) is an example of a first predicted box provided by the present application;
[0034] FIG. 5(b) is an example of a second predicted box provided by the present application;
[0035] FIG. 5(c) is an example of a third predicted box provided by the present application;
[0036] FIG. 5(d) is an example of a fourth predicted box provided by the present application;
[0037] FIG. 6(a) is an example of an alternative first predicted box provided by the present application;
[0038] FIG. 6(b) is an example of a third predicted box corresponding to the alternative first predicted box provided by the present application;
[0039] FIG. 6(c) is an example of a first target predicted box, an alternative first predicted box provided by the present application;
[0040] FIG. 6(d) is an example of a third predicted box corresponding to the first target predicted box, a third predicted box corresponding to the alternative first predicted box provided by the present application;
[0041] FIG. 6(e) is an example of an alternative second predicted box provided by the present application;
[0042] FIG. 6(f) is an example of a first predicted box corresponding to the alternative second predicted box provided by the present application;
[0043] FIG. 6(g) is an example of a second target predicted box, an alternative second predicted box provided by the present application;
[0044] FIG. 6(h) is an example of a first predicted box corresponding to the second target predicted box provided by the present application;
[0045] Figure 7 An example image of infrared labeled images provided for the present application;
[0046] Figure 8 A structural schematic diagram of a fault monitoring device of a power equipment provided for the present application. DETAILED DESCRIPTION
[0047] In the existing technical solutions for monitoring the fault state of power equipment in a power site based on images, different single-modal images have different advantages and disadvantages. Among them, different single-modal images can provide different kinds of information; for example, visible light images provide rich details when the light is sufficient, but the effect is poor at night or in low light conditions; infrared images are good at indicating the temperature of power equipment, suitable for night and bad weather, but lack of detailed information; although the fusion image combines the advantages of visible light images and infrared images, it is easy to produce information loss, color distortion and contrast balance problems in the fusion process, which may lead to inaccurate information.
[0048] For example, visible light images can provide detailed information of power equipment in a power site under sufficient light, but cannot provide detailed information of power equipment in a power site in the shadow; infrared images can provide temperature information of power equipment with abnormal surface temperature due to abnormal operation, but cannot provide additional information for power equipment with normal surface temperature or little temperature change; and the fusion image may neither accurately provide detailed information of power equipment in a power site under sufficient light nor provide temperature information of power equipment with abnormal surface temperature due to abnormal operation, because of the information loss, color distortion and contrast balance problems often occurring in the image fusion process. In summary, since the above various single-modal images can usually only provide part of the information useful for the fault monitoring of power equipment in the application of power equipment fault monitoring, it is possible to draw an inaccurate conclusion whether to choose any single-modal image to analyze whether the power equipment has failed.
[0049] To solve the above technical problems, the application provides a power equipment fault monitoring technical solution based on multi-modal images, which can break through the idea of "single-modal images can usually only provide part of the information useful for power equipment fault monitoring", and integrate and apply the part of useful information provided by different single-modal images to achieve "cross-image" application of "information useful for power equipment fault monitoring" in a novel way. To make the purpose, technical solutions and advantages of the embodiments of the application clearer, the technical solutions of the application will be described clearly and completely below in combination with the drawings. Obviously, the described embodiments are part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the application.
[0050] First, the application provides a power equipment fault monitoring method, as shown in Figure 1 Figure 1 The flowchart of the power equipment fault monitoring method provided by the application is shown, and the method comprises S110-S130, which are shown in detail as follows:
[0051] S110: detecting each single-modal image in the plurality of different types of single-modal images about the power site by a target detection model to obtain a single-modal labeled image.
[0052] In the single-modal labeled image, at least one prediction box and a confidence corresponding to each prediction box are labeled; the image content in the bounding box of the prediction box includes a target object recognized by the target detection model; and the confidence indicates the prediction probability of the target object being a power equipment.
[0053] Specifically, in the embodiments of the application, the image content displayed by the plurality of different types of single-modal images about the power site is the same or corresponding; for example, if the types of single-modal images include visible light images and infrared images, the above visible light images and infrared images should be images taken at the same shooting location with the same shooting angle and shooting distance. Since the visible light images and infrared images are taken at the same shooting location with the same shooting angle and shooting distance, the image content of the visible light images and the image content of the infrared images are corresponding.
[0054] It should be noted that in actual operation, the size of the visible light image is usually larger than that of the infrared image, and thus the image content displayed by the visible light image is usually more than that displayed by the infrared image, and thus if the type of the single-mode image of the power site is the visible light image and the infrared image, the image content displayed by the two single-mode images is "corresponding". In the embodiment of the present application, the target detection model is a neural network model for identifying and classifying target objects included in an image, such as a YOLO (You Only Look Once) model; wherein the target object is determined according to the requirement of training the YOLO model, and if the purpose of training the YOLO model is to identify the type of birds, the target object is various birds. In the embodiment of the present application, the target detection model used is used to identify and classify power equipment in the power site; wherein the type of power equipment that can be identified and classified is determined according to the training sample set for training the target detection model. In actual operation, when the plurality of single-mode images to be identified and classified are sequentially input into the target detection model, the target detection model sequentially outputs the identification and classification results of each single-mode image in the plurality of single-mode images.
[0055] The identification and classification result includes the position, confidence, type and type probability of the prediction box. Wherein, the prediction box is usually a rectangular box, and the prediction box is used to label the target object in the single-mode image in the form of boxing the target object, and thus the prediction box also has the function of positioning; in the embodiment of the present application, the target object refers to the power equipment; in actual operation, the target detection model outputs the center coordinates and the length and width of the prediction box When the center coordinates and the length and width of the prediction box are determined, the prediction box can be drawn in the single-mode image. The type refers to a certain "type of power equipment" to which the power equipment belongs. The confidence indicates the prediction probability that the target object is the power equipment; the type probability indicates the prediction probability that the power equipment belongs to a certain "type of power equipment"; both the confidence and the type probability are numbers between 0 and 1; the closer the confidence is to 1, the more clearly the target detection model thinks that the prediction box has labeled the power equipment; the closer the type probability is to 1, the more likely the target detection model thinks that the power equipment belongs to the corresponding type of power equipment. When the identification and classification result of each single-mode image is determined according to the target detection model, the single-mode labeled image corresponding to each single-mode image can be determined according to the identification and classification result.
[0056] In one implementation, the type of the single-mode image includes: a visible light image, a fusion image and an infrared augmented image; before S110, the method further includes steps (1) to (3), as shown below:
[0057] Step (1): Taking the to-be-labeled infrared image and the to-be-labeled visible light image about the power site at the same shooting location, at the same shooting angle and shooting distance.
[0058] Specifically, in the embodiment of the present application, the to-be-labeled visible light image and the to-be-labeled infrared image are images taken by a visible light camera and an infrared camera at the same shooting location, at the same shooting angle and shooting distance.
[0059] Step (2): Performing fusion processing on the to-be-labeled infrared image and the to-be-labeled visible light image to obtain a to-be-labeled fusion image.
[0060] Step (3): According to the size of the to-be-labeled infrared image and the size of the to-be-labeled visible light image, the image content of the to-be-labeled visible light image and the image content of the to-be-labeled infrared image, performing background filling at the edge of the to-be-labeled infrared image to obtain a to-be-labeled infrared augmented image.
[0061] Wherein, the size of the to-be-labeled infrared augmented image is the same as the size of the to-be-labeled visible light image; the relative position of the image content of the to-be-labeled infrared image in the to-be-labeled infrared augmented image is the same as the relative position of the image content of the to-be-labeled infrared image in the to-be-labeled visible light image.
[0062] Specifically, in actual operation, because the target detection model requires the size of the input image to be fixed when processing the image, if the sizes of the multiple to-be-labeled images are inconsistent, the connection weight cannot be matched, so that the output result cannot be calculated. Therefore, before inputting multiple single modal images such as the to-be-labeled infrared image, the to-be-labeled visible light image and the to-be-labeled fusion image into the target detection model, the multiple to-be-labeled images need to be adjusted to meet the size requirements of the target detection model for the input image; wherein the to-be-labeled infrared augmented image includes the image content of the to-be-labeled infrared image.
[0063] Generally, the size of the visible light image to be labeled among the different types of images to be labeled is the largest, so in order to ensure that the information of all images is not lost, the size of the visible light image to be labeled can be set as the size of the input image required by the target detection model, and before the infrared image to be labeled is input into the target detection model, the size of the infrared image to be labeled is first stretched to the same size as the visible light image. As shown in FIG. 2(a), FIG. 2(b), FIG. 2(c) and FIG. 2(d), FIG. 2(a) is an example of a visible light image to be labeled provided by the present application, which conforms to the size of the input image of the target detection model, FIG. 2(b) is an example of a fusion image to be labeled provided by the present application, which conforms to the size of the input image of the target detection model, FIG. 2(c) is an example of an infrared augmented image to be labeled provided by the present application, which conforms to the size of the input image of the target detection model, and FIG. 2(d) is an example of an infrared image to be labeled provided by the present application, which conforms to the size of the input image of the target detection model (i.e., the infrared image to be labeled after being stretched in size).
[0064] In the embodiments of the present application, the solid small square boxes and the circular boxes located in the large square box in FIG. 2(a) all represent target objects derived from the visible light image, the dashed small square boxes in FIG. 2(b), FIG. 2(c) and FIG. 2(d) represent target objects derived from the infrared image, and the circular box in FIG. 2(b) represents a target object derived from the visible light image; it should be emphasized that in order to briefly show the differences in image content between the visible light image to be labeled, the fusion image to be labeled, the infrared augmented image to be labeled and the infrared image to be labeled, the target objects are displayed in the form of prediction boxes in FIG. 2(a), FIG. 2(b), FIG. 2(c) and FIG. 2(d); in actual operation, the visible light image to be labeled, the fusion image to be labeled, the infrared augmented image to be labeled and the infrared image to be labeled do not contain prediction boxes. In actual operation, although stretching the infrared image to be labeled can make the stretched infrared image exhibit more detailed information, it may also cause information loss or even image distortion, so in order to preserve the un-stretched infrared image to be labeled, the present application constructs the infrared augmented image to be labeled, i.e., FIG. 2(c). The infrared image shown in FIG. 2(d) is the infrared image to be labeled after being stretched in size, and when the infrared image to be labeled shown in FIG. 2(d) is not stretched, the size of the dashed small square box is the same as that of the dashed small square box in FIG. 2(b).
[0065] As can be seen from FIG. 2(a), FIG. 2(c) and FIG. 2(d), the size of the infrared augmented image to be labeled shown in FIG. 2(c) is the same as that of the visible light image to be labeled shown in FIG. 2(a), and the infrared augmented image to be labeled shown in FIG. 2(c) includes the image content of the un-stretched infrared image to be labeled (i.e., the image content shown by the dashed small square box).
[0066] S120: According to the confidence, determine a target bounding box from at least one bounding box in the plurality of single-modal annotated images for annotating the same target object.
[0067] Specifically, in order to solve the technical problem that any single-modal image may not be accurate enough to analyze whether the power equipment fails, the information useful for monitoring the failure of the power equipment in the single-modal annotated image corresponding to each single-modal image is extracted from the image, and then is comprehensively applied, i.e., the comprehensive application across images. The image content in each bounding box in the single-modal annotated image is the information useful for monitoring the failure of the power equipment in the corresponding single-modal image, because the image content in the bounding box includes the power equipment to be monitored for failure. In actual operation, although the image content in the bounding box of each bounding box includes the power equipment, i.e., the image content in the bounding box belongs to the information useful for monitoring the failure of the power equipment, only the bounding box with the highest useful degree of the information useful for monitoring the failure of the power equipment can correspond to the most accurate monitoring result of the failure.
[0068] In the embodiment of the present application, the confidence of the bounding box is used as the measurement standard of the useful degree, i.e., for the power equipment in the bounding box in different single-modal annotated images, the higher the confidence of the bounding box in which the power equipment is located, the higher the useful degree of the information useful for monitoring the failure of the power equipment provided by the bounding box. Therefore, the bounding box with the highest confidence from at least one bounding box in the plurality of single-modal annotated images for annotating the same target object is determined as the target bounding box of the corresponding target object. In actual operation, the number of the bounding box with the highest confidence can be one or multiple, i.e., there can be multiple bounding boxes with the same and highest confidence. For example, the number of the bounding boxes for annotating the same target object is four, and the corresponding confidences are 0.85, 0.74, 0.84 and 0.85, respectively, wherein the confidence 0.85 is the highest confidence and the bounding box corresponding to the confidence 0.85 has two. In actual operation, when there are multiple bounding boxes with the same and highest confidence, one target bounding box can be selected from the multiple bounding boxes with the highest confidence, and the selection standard can be determined according to actual needs, which is not limited in the present application.
[0069] The target bounding box is the bounding box that can provide the information useful for monitoring the failure of the power equipment with the highest useful degree for the image content in the bounding box.
[0070] In an implementation manner, S120 comprises steps (4) to (6), details of which are shown as follows:
[0071] Step (4): Select a reference annotation image from the plurality of single-modal annotation images, and determine all the prediction boxes in the reference annotation image as reference prediction boxes.
[0072] Specifically, according to the actual situation, it is known that each single-modal image used in the embodiment of the application is an image of a power site, and the power site usually includes a plurality of power equipment. Therefore, each single-modal annotation image generally includes at least one prediction box, and the plurality of single-modal annotation images will include a large number of prediction boxes. Therefore, it is difficult to select the prediction boxes for labeling the same target object from the large number of prediction boxes.
[0073] In actual operation, if the image content included in the bounding boxes of different prediction boxes in the single-modal annotation image is different types of power equipment, the prediction boxes for labeling "the same power equipment" can be directly selected according to the "type" output by the target detection model. However, in actual application, the image content included in the bounding boxes of different prediction boxes in the single-modal annotation image is usually the same type of power equipment. Therefore, it is not possible to accurately determine the prediction boxes for labeling "the same power equipment" in any two single-modal annotation images only according to the "type" output by the target detection model. In the embodiment of the application, whether two prediction boxes are for labeling "the same power equipment" is determined by determining the similarity between the image content in the bounding boxes of the two prediction boxes. If the similarity between the image content in the bounding boxes of the two prediction boxes is greater than or equal to a preset similarity threshold, it is considered that the two prediction boxes are for labeling "the same power equipment". If the similarity between the image content in the bounding boxes of the two prediction boxes is less than the similarity threshold, it is considered that the two prediction boxes are not for labeling "the same power equipment". However, in actual operation, if all the prediction boxes included in all the single-modal images are to be confirmed for similarity, the workload will be very large.
[0074] In order to improve the efficiency of determining the prediction boxes for labeling "the same power equipment", the embodiment of the application determines one single-modal annotation image from the plurality of single-modal annotation images as a reference annotation image, and calculates the similarity between the image content in the bounding boxes of each prediction box in the single-modal annotation image other than the reference annotation image and the image content in the bounding boxes of the prediction boxes in the single-modal image, to determine the prediction boxes for labeling "the same power equipment" in the plurality of single-modal annotation images.
[0075] In the embodiments of the present application, the reference annotation image is determined according to actual requirements. For example, if the current fault monitoring relies more on the surface details of the power equipment to determine whether a fault occurs, the visible light annotation image can be determined as the reference annotation image. If the current fault monitoring relies more on the temperature information of the power equipment to determine whether a fault occurs, the infrared annotation image can be determined as the reference annotation image.
[0076] When the reference annotation image is determined, the prediction box in the reference annotation image is the reference prediction box.
[0077] Step (5): Among all the prediction boxes included in the single-modality annotation images remaining after the reference annotation image is removed from the plurality of single-modality annotation images, the prediction box whose image content in the bounding box is greater than or equal to the similarity threshold value between the image content in the bounding box of the reference prediction box is determined as a similar prediction box.
[0078] Specifically, the similar prediction box is a prediction box determined by the similarity threshold value in the single-modality annotation images remaining after the reference annotation image is removed from the plurality of single-modality annotation images, and the prediction box is used to annotate the same power equipment as the reference prediction box. In actual application, the number of similar prediction boxes can be 0. When the number of similar prediction boxes is 0, it indicates that there is a target object in the reference annotation image that does not exist in other single-modality images. For example, if the reference annotation image is a visible light annotation image, since the image size of the visible light annotation image is generally larger than that of the infrared annotation image, the visible light annotation image can include more image content than the infrared annotation image. Therefore, the visible light annotation image can include more prediction boxes than the infrared annotation image. Therefore, when the visible light annotation image is the reference annotation image, the infrared annotation image can not have a similar prediction box corresponding to a reference prediction box in the visible light annotation image.
[0079] Step (6): For each reference prediction box, the prediction box with the highest confidence in the reference prediction box and the corresponding similar prediction box is determined as the target prediction box of the target object corresponding to the reference prediction box.
[0080] Specifically, after the similar prediction boxes and the reference prediction boxes used to annotate the same power equipment in the plurality of single-modality annotation images are determined, the prediction box with the highest useful degree of information useful for fault monitoring of the power equipment can be determined according to the confidence, that is, the target prediction box of the power equipment. If a reference prediction box does not have a corresponding similar prediction box, the reference prediction box can be determined as a target prediction box.
[0081] In an implementation manner, the types of the single-modal labeled images include: visible light labeled images, fusion labeled images, and infrared augmented labeled images; S120 comprises steps (7) to (14), which are shown as follows:
[0082] Step (7): according to the image content of the infrared image to be labeled, the image region corresponding to the image content of the infrared image to be labeled in the visible light labeled image and the fusion labeled image is determined as a visible light labeled region and a fusion labeled region respectively.
[0083] Specifically, as shown in FIG. 3(a), FIG. 3(b) and FIG. 3(c), FIG. 3(a) is an example of the visible light labeled image provided by the present application, FIG. 3(b) is an example of the fusion labeled image provided by the present application, and FIG. 3(c) is an example of the infrared augmented labeled image provided by the present application. Since the image contents of different single-modal labeled images are different, the prediction boxes in the corresponding single-modal labeled images are also not completely the same.
[0084] In the embodiment of the present application, the focus of fault monitoring is placed on the power equipment capable of simultaneously determining the detail information and the temperature information, that is, only the multiple “same power equipment” appearing in the infrared image and the visible light image are monitored for fault, and therefore the target prediction box to be determined is also the target prediction box of the multiple “same power equipment”.
[0085] In actual operation, in order to ensure that the power equipment corresponding to the determined target prediction box is the power equipment appearing in the infrared labeled image and the visible light labeled image at the same time, the visible light labeled region and the fusion labeled region are intercepted, so that the image contents of the visible light labeled region and the fusion labeled region correspond to the image content of the infrared labeled image. As shown in FIG. 4(a) and FIG. 4(b), FIG. 4(a) is an example of the visible light labeled region provided by the present application, and FIG. 4(b) is an example of the fusion labeled region provided by the present application. The image content in the frame of the rectangular frame located in the central region of the image in FIG. 4(a) and FIG. 4(b) is the visible light labeled region and the fusion labeled region. It should be emphasized that in the process of monitoring the power equipment in the power site for fault, many images will be taken. For the power equipment for which the target prediction box is not determined in this fault monitoring, the corresponding target prediction box can be determined in the fault monitoring based on other images, and then the fault prediction is performed accordingly.
[0086] Step (8): the prediction boxes included in the visible light labeled region are all determined as first prediction boxes, the prediction boxes included in the fusion labeled region are all determined as second prediction boxes, and the prediction boxes included in the infrared augmented labeled image are all determined as third prediction boxes.
[0087] Specifically, as shown in Table 1, Table 1 is a detailed table of the prediction boxes in each single-modal labeled image / labeled region provided by the present application; wherein the first prediction box has the meaning of the i-th first prediction box in the first prediction boxes included in the visible light labeled region ; the second prediction box has the meaning of the i-th second prediction box in the second prediction boxes included in the fusion labeled region ; the third prediction box has the meaning of the i-th third prediction box in the third prediction boxes included in the infrared augmented labeled image , , , , and , , , are all positive integers. Table 1 Detailed table of prediction boxes in each single-modal labeled image / labeled region
[0088]
[0089]
[0090] It is emphasized that the first prediction box and the second prediction box determined in the embodiment of the present application are both "complete" prediction boxes in the visible light marked area and the fusion marked area, because the visible light marked area and the fusion marked area are image areas obtained by cutting in the visible light marked image and the fusion marked image, and only include part of the image content in the visible light marked image and the fusion marked image, and usually only include part of the prediction boxes in the visible light marked image and the fusion marked image. According to the aforementioned "therefore, it is necessary to cut the visible light marked image and the fusion marked image, so that the image content of the cut visible light marked area and the fusion marked area corresponds to the image content of the infrared marked image", it can be known that the cutting basis of the visible light marked area and the fusion marked area obtained in the embodiment of the present application is "the image content of the obtained marked image is the same as the image content of the infrared marked image", and because the positions of the prediction boxes in the visible light marked image and the fusion marked image may appear at any position in the image in actual operation, in some cases, there may be a cut-off prediction box in the visible light marked area and the fusion marked area obtained by cutting in the visible light marked image and the fusion marked image. In the embodiment of the present application, because the fault monitoring is performed based on the similarity between the prediction boxes in different marked images and / or different marked areas, therefore the prediction boxes used for comparing the similarity must have complete "image content in the prediction box", if the "cut-off prediction box" is used to calculate the similarity, the calculated similarity will be very low, which will seriously affect the determination of the fault monitoring result, therefore in the process of determining the first prediction box and the second prediction box, the selected are all "complete" prediction boxes.
[0091] Step (9): for each third prediction box, the first prediction box in which the similarity between the image content in the frame and the image content in the frame of the third prediction box is greater than or equal to the similarity threshold is determined as the candidate first prediction box.
[0092] Specifically, in the embodiments of this application, as shown in Figures 5(a), 5(b), 5(c), and 5(d), Figure 5(a) is an example diagram of the first prediction box provided by this application, Figure 5(b) is an example diagram of the second prediction box provided by this application, Figure 5(c) is an example diagram of the third prediction box provided by this application, and Figure 5(d) is an example diagram of the fourth prediction box provided by this application. In Figure 5(a), the image area within the large solid-line box represents the visible light labeled area. The small solid-line boxes within the large solid-line box in Figure 5(a) represent prediction boxes originating from the visible light image and located within the visible light labeled area. The circular boxes within the large box in Figure 5(a) represent prediction boxes originating from the visible light image and located outside the visible light labeled area. In Figure 5(b), the image area within the large dashed-line box represents the fusion labeled area. The small dashed-line boxes in Figure 5(b) represent prediction boxes for the image portion that has fused visible light and infrared image information (for visual representation of fusion, the small dashed-line box of the infrared image prediction box is still used), i.e., prediction boxes located within the fusion labeled area. The circular boxes in Figure 5(b) represent prediction boxes originating from the visible light image and located outside the fusion labeled area. The information filled in the prediction boxes in Figures 5(a), 5(b), 5(c), and 5(d) is... , , , The corresponding values represent the confidence level of the prediction box.
[0093] In practice, due to infrared amplification and annotation of images The image content includes unstretched infrared images; therefore, in determining the target prediction bounding box, the visible light-marked area is first targeted. The first predicted bounding box and infrared amplified labeled image in the image. The third prediction box undergoes the first round of voting to select the prediction boxes used to label "the same power equipment".
[0094] ;
[0095] In the formula, Indicates the visible light marked area Included The first prediction box The first prediction box; Indicates infrared amplified labeled image Included The third prediction box A third prediction box; Indicates an image labeled with infrared amplification. The included third prediction box Voting is conducted primarily on the selected images, meaning the screening process involves using infrared amplified and labeled images. comprise a third prediction box as a reference; representing the visible light annotation region and the infrared augmented annotation image voted by two single-modality annotation images / annotation regions.
[0096] In actual operation, for each of the third prediction boxes in the infrared augmented annotation image comprise , determine the first prediction boxes in which the similarity between the image content within the bounding box and the image content within the bounding box of the third prediction box is greater than or equal to the similarity threshold value. According to the above discussion, if the similarity between the image content within the bounding box of the two prediction boxes is greater than or equal to the similarity threshold value, it is considered that the two prediction boxes are used to annotate the same power equipment. As shown in FIG. 6(a) and FIG. 6(b), FIG. 6(a) is an example diagram of the alternative first prediction boxes provided by the present application, and FIG. 6(b) is an example diagram of the third prediction boxes corresponding to the alternative first prediction boxes provided by the present application. The alternative first prediction boxes obtained after step (9) are annotated in FIG. 6(a), and the third prediction boxes corresponding to the alternative first prediction boxes obtained after step (9) are annotated in FIG. 6(b).
[0097] In the embodiments of the present application, the first prediction boxes in the visible light annotation region and the third prediction boxes in the infrared augmented annotation image used to annotate the same power equipment are determined as the alternative first prediction boxes. It should be emphasized that the number of alternative first prediction boxes can be multiple, because the image content of the visible light region image and the image content of the infrared augmented annotation image can include multiple same power equipment, and therefore the corresponding visible light annotation region and the infrared augmented annotation image may include multiple groups of alternative first prediction boxes and third prediction boxes used to annotate the same power equipment, and each group of alternative first prediction boxes and third prediction boxes corresponds to one same power equipment.
[0098] Step (10): For each third prediction box, the prediction box with the highest confidence in the third prediction box and the corresponding alternative first prediction box is determined as the first target prediction box.
[0099] Specifically, after determining multiple groups of alternative first prediction boxes and third prediction boxes used to annotate the same power equipment, the prediction box with the highest useful degree of the information useful for the fault monitoring of the power equipment can be selected as the first target prediction box from each group of alternative first prediction boxes and third prediction boxes according to the confidence.
[0100] In actual operation, the third prediction frame remaining in the infrared augmented annotation image after the first target prediction frame is determined is determined as a candidate third prediction frame for subsequent voting.
[0101] As shown in Table 2, Table 2 is a detailed table of prediction frames in each single modality annotation image / annotation region after the first round of voting provided by the present application, wherein the details of the candidate first prediction frame, the candidate first prediction frame and the like are mainly written.
[0102] Table 2: Detailed table of prediction frames in each single modality annotation image / annotation region
[0103]
[0104] As shown in FIG. 6(c) and FIG. 6(d), FIG. 6(c) is an example diagram of the first target prediction frame, the candidate first prediction frame provided by the present application, and FIG. 6(d) is an example diagram of the third prediction frame corresponding to the first target prediction frame, the third prediction frame corresponding to the candidate first prediction frame provided by the present application; in FIG. 6(c), the first target prediction frame obtained after step (10), the candidate first prediction frame remaining after removing the first target prediction frame, and the candidate first prediction frame are annotated, and the candidate first prediction frame is the prediction frame remaining in the visible light annotation region after removing the candidate first prediction frame and the first target prediction frame; in FIG. 6(d), the third prediction frame corresponding to the candidate first prediction frame, the third prediction frame corresponding to the remaining candidate first prediction frame, and the third prediction frame corresponding to the candidate first prediction frame are annotated after step (10).
[0105] It should be emphasized that although in the example shown in FIG. 6(c), the first target prediction frame is derived from the first prediction frame, in the embodiments of the present application, since the determined first target prediction frame is determined in the third prediction frame and the candidate first prediction frame corresponding to the third prediction frame, there is a case of determining the third prediction frame as the first target prediction frame; if there is a case of determining the third prediction frame as the first target prediction frame, it means that in the current fault monitoring process, the image content in the frame of the third prediction frame in the infrared augmented annotation image is clearer, which is also the goal of the technical solution of the embodiments of the present application, that is, on the premise of acknowledging the existence of various image acquisition errors and the like operation errors, the prediction frame that can most clearly represent the power equipment is determined, so as to determine the most accurate fault monitoring result according to the prediction frame.
[0106] Step (11): determining, as a candidate second prediction box, a second prediction box in all the second prediction boxes whose similarity between the image content within the bounding box and the image content within the bounding box of the candidate first prediction box is greater than or equal to the similarity threshold. The candidate first prediction box is a prediction box remaining in the visible light annotation region after removing the candidate first prediction box and the first target prediction box (if the first target prediction box exists in the visible light annotation region).
[0107] Specifically, in the embodiment of the present application, after the first round of voting, the second round of voting is performed on the candidate first prediction boxes and the second prediction boxes in the visible light annotation region and the fusion annotation region , for screening the prediction boxes for labeling "the same power equipment".
[0108] ;
[0109] In the formula, represents the i-th candidate first prediction box in the visible light annotation region ; represents the j-th second prediction box in the fusion annotation region . In actual operation, for each candidate first prediction box in the visible light annotation region , the second prediction boxes whose similarity between the image content within the bounding box and the image content within the bounding box of the candidate first prediction box is greater than or equal to the similarity threshold are determined. In the embodiment of the present application, the second prediction box in the fusion annotation region , which is used for labeling "the same power equipment" with the candidate first prediction box in the visible light annotation region , is determined as a candidate second prediction box. As shown in FIG. 6 (e) and FIG. 6 (f), FIG. 6 (e) is an example diagram of the candidate second prediction box provided by the present application, and FIG. 6 (f) is an example diagram of the first prediction box corresponding to the candidate second prediction box provided by the present application; the candidate second prediction box obtained after step (11) is labeled in FIG. 6 (e), and the first prediction box corresponding to the candidate second prediction box, the candidate first prediction box, the first target prediction box and the remaining candidate first prediction box obtained after step (11) are labeled in FIG. 6 (f).
[0110] Step (12): determining, as a second target prediction box, the prediction box with the maximum confidence between the candidate first prediction box and the candidate second prediction box for each candidate first prediction box.
[0111] Step (12): determining, as a second target prediction box, the prediction box with the maximum confidence between the candidate first prediction box and the candidate second prediction box for each candidate first prediction box.
[0112] Specifically, after determining the multiple groups of candidate first prediction boxes and candidate second prediction boxes for labeling "the same power equipment", in each group of candidate first prediction boxes and candidate second prediction boxes, the prediction box with the highest "usefulness" of "information useful for failure monitoring of the power equipment" can be selected as the second target prediction box according to the confidence.
[0113] As shown in Table 3, Table 3 is a detailed table of prediction boxes in each single-modal labeled image / region provided by the present application after the second round of voting, which mainly shows the details of the candidate first prediction boxes, candidate first prediction boxes and the like.
[0114] Table 3 is a detailed table of prediction boxes in each single-modal labeled image / region
[0115]
[0116] As shown in FIG. 6(g) and FIG. 6(h), FIG. 6(g) is an example diagram of the second target prediction box and the candidate second prediction box provided by the present application, and the candidate second prediction box is the second prediction box remaining after removing the candidate second prediction box and the second target prediction box in the fusion labeled region; FIG. 6(h) is an example diagram of the first prediction box corresponding to the second target prediction box provided by the present application; the second target prediction box and the candidate second prediction box in FIG. 6(g) are labeled after step (12), and the first prediction box corresponding to the candidate second prediction box, the first prediction box corresponding to the second target prediction box, the remaining candidate first prediction box, the candidate first prediction box and the first target prediction box in FIG. 6(h) are labeled after step (12).
[0117] It should be emphasized that although in the example shown in FIG. 6(g), the "second target prediction box" is derived from the candidate second prediction box, in the embodiments of the present application, since the determined second target prediction box is in the "candidate first prediction box" and the "candidate second prediction box", there is a case of determining the "candidate second prediction box" as the second target prediction box.
[0118] Step (13): for each candidate third prediction box, the second prediction box in all candidate second prediction boxes, the similarity between the image content in the bounding box and the image content in the bounding box of each candidate third prediction box is less than the similarity threshold, is determined as the third target prediction box.
[0119] Among them, the candidate second prediction box is the second prediction box remaining after removing the candidate second prediction box and the second target prediction box (if the second target prediction box exists in the fusion labeled region) in the fusion labeled region; the candidate third prediction box is the remaining third prediction box which is not determined as the first target prediction box.
[0120] Specifically, in this embodiment of the application, after the second round of voting, the fusion annotation region is now... and infrared amplified labeled images The candidate second and candidate third prediction boxes are voted on in a third round to select the prediction box used to label "the same power equipment".
[0121] ;
[0122] In the formula, Indicates the merged annotation area The first in One candidate second prediction box; Indicates infrared amplified labeled image The first in There are several candidate third prediction boxes. In practice, this is applied to infrared amplified and labeled images. Among the candidate third prediction boxes, determine the candidate second prediction boxes whose similarity between the image content inside the border of each candidate second prediction box and the image content inside the border of each candidate third prediction box is less than the similarity threshold.
[0123] In this embodiment of the application, the fused annotation region will be... Infrared amplified labeled images The candidate third prediction box in the image is used to label the candidate second prediction box for "the same power equipment" and is determined as the third target prediction box. In this embodiment, the image information in the fused annotation region of the fused annotation image and the image information in the infrared amplified annotation image both contain image information from the infrared image. Therefore, the similarity between the candidate second prediction box and the candidate third prediction box corresponding to the same target object in the fused annotation region and the infrared amplified annotation image must be very high (basically 100%). Therefore, the selection principle for determining the third target prediction box is not "similarity greater than or equal to the similarity threshold".
[0124] It is emphasized that the reason why the similarity between the second candidate prediction box and the third candidate prediction box mentioned above must be very high is that the similarity between the second candidate prediction box and the third candidate prediction box corresponding to the same target object is very high. However, as shown in FIG. 6(d) and FIG. 6(g), the remaining second candidate prediction boxes and third candidate prediction boxes are not all corresponding to the same target object, so the similarity between the remaining second candidate prediction boxes and third candidate prediction boxes is not very high. Only the similarity between the second candidate prediction box and the third candidate prediction box corresponding to the same target object is very high. Therefore, the role of step (13) is to collect the remaining second candidate prediction boxes and third candidate prediction boxes in the above comparison process to jointly vote in the next round. Therefore, the screening principle for determining the third target prediction box is "less than the similarity threshold value", rather than "greater than the similarity threshold value", because the second prediction box or the third prediction box greater than the similarity threshold value has been determined as the second target prediction box and the third target prediction box, respectively.
[0125] Step (14): determining the first target prediction box, the second target prediction box and the third target prediction box as the target prediction box.
[0126] Specifically, after three rounds of voting, the first target prediction box, the second target prediction box and the third target prediction box used for labeling "the same power equipment" in the visible light labeling region , the fusion labeling region and the infrared augmented labeling image can be determined, and then the target prediction box can be obtained by integrating them. In the embodiment of the present application, the prediction boxes in the visible light labeling region , the fusion labeling region and the infrared augmented labeling image can label the power equipment in the image content of the corresponding visible light image, fusion image and infrared augmented image, so that the accurate target detection box can be determined according to the visible light labeling region , the fusion labeling region and the infrared augmented labeling image .
[0127] In an implementation manner, the type of the single-modal labeling image further includes an infrared labeling image. After step (14), the method further includes steps (15) and (16), as shown below:
[0128] Step (15): determining the prediction box included in the infrared labeling image as a fourth prediction box.
[0129] Specifically, as shown in Figure 7 ,Figure 7 The fourth prediction box on the example infrared labeled image provided in the present application is a prediction box.
[0130] Step (16): supplementing, to the target prediction box, the fourth prediction box in which the similarity between the image content in the bounding box and the image content in the bounding box of the target prediction box is less than the similarity threshold.
[0131] Specifically, in the embodiments of the present application, the fourth prediction box of the infrared labeled image mainly plays a role of "checking for omissions and making up for omissions", because the infrared labeled image in the embodiments of the present application is a single-modal labeled image corresponding to the stretched infrared image, which can include prediction boxes of power equipment that are not labeled in the visible light labeled region , the fusion labeled region and the infrared augmented labeled image Therefore, the fourth prediction box remaining after removing the prediction box of the infrared labeled image that is used to label "the same power equipment" as the target prediction box is a prediction box unexpectedly obtained by the infrared image due to stretching.
[0132] S130: determining the fault monitoring result of the corresponding power equipment according to the image content in the bounding box of the target prediction box.
[0133] Specifically, after the target prediction box is determined, the subsequent fault monitoring result of the corresponding power equipment can be determined according to the image content in the bounding box of the target prediction box; the technical solution for determining whether the corresponding power equipment has a fault according to the image content in the bounding box can be determined according to actual conditions, which is not limited in the present application.
[0134] As can be seen from the above, the embodiments of the present application take the prediction box for labeling the power equipment as a carrier of "information useful for fault monitoring of the power equipment", and take the confidence of the prediction box as a measurement standard of "usefulness" of "information useful for fault monitoring of the power equipment", for determining the most useful target prediction box for labeling the same target object from multiple single-modal labeled images, and subsequently performing fault diagnosis on the power equipment according to the image content in the bounding box of the target prediction box.
[0135] Secondly, the present application provides a power equipment fault monitoring device, as shown in Figure 8 . Figure 8A structural schematic diagram of a fault monitoring device of a power equipment is provided in the present application, and the device comprises: a target detection module 310, a target screening module 320, and a fault diagnosis module 330. The target detection module 310 is configured to detect each of a plurality of different types of single-modal images about a power site by a target detection model to obtain a single-modal labeled image. The single-modal labeled image is labeled with at least one prediction box and a confidence corresponding to each prediction box. The image content in the edge box of the prediction box includes a target object recognized by the target detection model. The confidence indicates a prediction probability that the target object is a power equipment. The target screening module 320 is configured to determine a target prediction box in at least one prediction box for labeling the same target object in a plurality of single-modal labeled images according to the confidence. The fault diagnosis module 330 is configured to determine a fault monitoring result of the corresponding power equipment according to the image content in the edge box of the target prediction box.
[0136] In an implementation manner, the target screening module 320 is further configured to select a reference labeled image from the plurality of single-modal labeled images, and determine all prediction boxes in the reference labeled image as reference prediction boxes. The target screening module 320 is further configured to determine a similar prediction box from all prediction boxes in the single-modal labeled images remaining after removing the reference labeled image, wherein the image content in the edge box of the similar prediction box is similar to the image content in the edge box of the reference prediction box with a similarity greater than or equal to a similarity threshold. The target screening module 320 is further configured to determine, for each reference prediction box, a prediction box with the maximum confidence from the reference prediction box and the corresponding similar prediction box as the target prediction box of the target object corresponding to the reference prediction box.
[0137] In an implementation manner, the types of the single-modal images include visible light images, fusion images, and infrared augmented images. The device further comprises a single-modal image module. The single-modal image module is configured to capture a to-be-labeled infrared image and a to-be-labeled visible light image about a power site at the same shooting location, the same shooting angle, and the same shooting distance. The single-modal image module is further configured to perform fusion processing on the to-be-labeled infrared image and the to-be-labeled visible light image to obtain a to-be-labeled fusion image. The single-modal image module is further configured to perform background padding at an edge of the to-be-labeled infrared image according to the size of the to-be-labeled infrared image and the size of the to-be-labeled visible light image, the image content of the to-be-labeled visible light image, and the image content of the to-be-labeled infrared image to obtain a to-be-labeled infrared augmented image. The size of the to-be-labeled infrared augmented image is the same as the size of the to-be-labeled visible light image. The relative position of the image content of the to-be-labeled infrared image in the to-be-labeled infrared augmented image is the same as the relative position of the image content of the to-be-labeled infrared image in the to-be-labeled visible light image.
[0138] In an implementation manner, the types of the single-modal labeled images include: a visible light labeled image, a fusion labeled image, and an infrared augmented labeled image; the target screening module 320 is further configured to determine, according to image content of the infrared image to be labeled, image regions corresponding to the image content of the infrared image to be labeled in the visible light labeled image and the fusion labeled image as visible light labeled regions and fusion labeled regions respectively; the target screening module 320 is further configured to determine all the bounding boxes included in the visible light labeled regions as first bounding boxes, determine all the bounding boxes included in the fusion labeled regions as second bounding boxes, and determine all the bounding boxes included in the infrared augmented labeled image as third bounding boxes; the target screening module 320 is further configured to, for each third bounding box, determine, from all the first bounding boxes, a first bounding box in which a similarity between image content in a bounding box and image content in a bounding box of the third bounding box is greater than or equal to a similarity threshold value, as a candidate first bounding box; and the target screening module 320 is further configured to, for each third bounding box, determine, from the third bounding box and the corresponding candidate first bounding box, a bounding box with a maximum confidence as a first target bounding box.
[0139] In an implementation manner, the target screening module 320 is further configured to determine, from all the second bounding boxes, a second bounding box in which a similarity between image content in a bounding box and image content in a bounding box of a candidate first bounding box is greater than or equal to a similarity threshold value, as a candidate second bounding box; the candidate first bounding box is a bounding box remaining in the visible light labeled regions after the candidate first bounding box and the first target bounding box are removed; the target screening module 320 is further configured to, for each candidate first bounding box, determine, from the candidate first bounding box and a bounding box with a maximum confidence of the candidate second bounding box, a bounding box as a second target bounding box; the target screening module 320 is further configured to, for each candidate third bounding box, determine, from all the candidate second bounding boxes, a second bounding box in which a similarity between image content in a bounding box and image content in a bounding box of each candidate third bounding box is less than a similarity threshold value, as a third target bounding box; the candidate second bounding box is a second bounding box remaining in the fusion labeled regions after the candidate second bounding box and the second target bounding box are removed; the candidate third bounding box is a remaining third bounding box that is not determined as the first target bounding box; and the target screening module 320 is further configured to determine the first target bounding box, the second target bounding box, and the third target bounding box as target bounding boxes.
[0140] In an implementation manner, the types of the single-modal labeled images further include an infrared labeled image; the target screening module 320 is further configured to determine bounding boxes included in the infrared labeled image as fourth bounding boxes; and the target screening module 320 is further configured to, from all the fourth bounding boxes, determine a fourth bounding box in which a similarity between image content in a bounding box and image content in a bounding box of a target bounding box is less than a similarity threshold value, and supplement the fourth bounding box to the target bounding box.
[0141] Thirdly, the present application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of S110-S130 provided by the above-mentioned embodiments.
[0142] Fourthly, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable on the processor to execute the steps of S110-S130 of the above-mentioned embodiments.
[0143] Fifthly, the computer program product provided by the present application comprises a computer readable storage medium storing program codes, and the instructions included in the program codes are used to execute the method in the above-mentioned method embodiments. The specific implementation can be referred to the steps of S110-S130 of the method embodiments, and will not be described here.
[0144] In the embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There can be another division during actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some communication interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0145] In addition, the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the present embodiment.
Claims
1. A method of fault monitoring of an electrical power device, characterized by, The method comprises: detecting each of a plurality of different types of single modal images about the power site by a target detection model to obtain single modal labeled images; the single modal images comprise visible light images, fusion images and infrared augmented images; the single modal labeled images comprise visible light labeled images, fusion labeled images and infrared augmented labeled images; wherein at least one prediction box and a confidence corresponding to each prediction box are labeled on the single modal labeled images; the image content in the edge box of the prediction box comprises a target object recognized by the target detection model; the confidence indicates the prediction probability that the target object is a power equipment; performing fusion processing on the to-be-labeled infrared image and the to-be-labeled visible light image about the power site to obtain a to-be-labeled fusion image; performing background padding at the edge of the to-be-labeled infrared image to obtain a to-be-labeled infrared augmented image according to the size of the to-be-labeled infrared image and the size of the to-be-labeled visible light image, the image content of the to-be-labeled visible light image and the image content of the to-be-labeled infrared image; wherein the size of the to-be-labeled infrared augmented image is the same as the size of the to-be-labeled visible light image; the relative position of the image content of the to-be-labeled infrared image in the to-be-labeled infrared augmented image is the same as the relative position of the image content of the to-be-labeled infrared image in the to-be-labeled visible light image; determining a target prediction box in at least one prediction box for labeling the same target object in a plurality of single modal labeled images according to the confidence; determining a fault monitoring result of the corresponding power equipment according to the image content in the edge box of the target prediction box; determining a target prediction box in at least one prediction box for labeling the same target object in a plurality of single modal labeled images according to the confidence, comprising: determining an image region corresponding to the image content of the to-be-labeled infrared image in the visible light labeled image and the fusion labeled image as a visible light labeled region and a fusion labeled region respectively according to the image content of the to-be-labeled infrared image; determining all prediction boxes included in the visible light labeled region as first prediction boxes, determining all prediction boxes included in the fusion labeled region as second prediction boxes, and determining all prediction boxes included in the infrared augmented labeled image as third prediction boxes; for each third prediction box, determining a first prediction box in all first prediction boxes, in which the similarity between the image content in the edge box of the first prediction box and the image content in the edge box of the third prediction box is greater than or equal to a similarity threshold, as a candidate first prediction box; for each third prediction box, determining a prediction box with the maximum confidence in the third prediction box and the corresponding candidate first prediction box as a first target prediction box; determining a second prediction box in all second prediction boxes, in which the similarity between the image content in the edge box of the second prediction box and the image content in the edge box of the candidate first prediction box is greater than or equal to the similarity threshold, as a candidate second prediction box; The candidate first prediction frame is a prediction frame remaining after the visible light annotation region is removed from the candidate first prediction frame and the first target prediction frame. For each candidate first prediction frame, the candidate first prediction frame and the prediction frame with the highest confidence in the candidate second prediction frame are determined as a second target prediction frame. For each candidate third prediction frame, the second prediction frame with a similarity between image content in a bounding box of the second prediction frame and image content in a bounding box of each candidate third prediction frame being less than the similarity threshold is determined as a third target prediction frame. The candidate second prediction frame is a second prediction frame remaining after the fusion annotation region is removed from the candidate second prediction frame and the second target prediction frame; and the candidate third prediction frame is a remaining third prediction frame that is not determined as the first target prediction frame. The first target prediction frame, the second target prediction frame, and the third target prediction frame are determined as the target prediction frame.
2. The method of claim 1, wherein, The method further comprises: selecting a reference annotation image from the plurality of single-modal annotation images, and determining all prediction frames in the reference annotation image as reference prediction frames; determining, from all prediction frames in the single-modal annotation images remaining after the reference annotation image is removed, a prediction frame with a similarity between image content in a bounding box of the prediction frame and image content in a bounding box of the reference prediction frame being greater than or equal to a similarity threshold as a similar prediction frame; for each reference prediction frame, determining a prediction frame with the highest confidence from the reference prediction frame and the corresponding similar prediction frame as a target prediction frame of a target object corresponding to the reference prediction frame.
3. The method of claim 1, wherein, The method further comprises: capturing a to-be-annotated infrared image and a to-be-annotated visible light image about the power site at the same shooting location, the same shooting angle, and the same shooting distance.
4. The method of claim 1, wherein, The single-modal annotation image further comprises an infrared annotation image; and after the first target prediction frame, the second target prediction frame, and the third target prediction frame are determined as the target prediction frame, the method further comprises: determining prediction frames included in the infrared annotation image as fourth prediction frames; complementing, from all fourth prediction frames, the fourth prediction frame with a similarity between image content in a bounding box of the fourth prediction frame and image content in a bounding box of the target prediction frame being less than the similarity threshold to the target prediction frame.
5. A failure monitoring device of a power equipment characterized by comprising: The apparatus comprises a target detection module, a target screening module, and a fault diagnosis module. The target detection module is configured to detect each single-modal image in a plurality of single-modal images of different types about a power site by a target detection model to obtain a single-modal annotation image; the single-modal image comprises a visible light image, a fusion image, and an infrared amplification image; and the single-modal annotation image comprises a visible light annotation image, a fusion annotation image, and an infrared amplification annotation image. The single-modal labeled image is labeled with at least one prediction box and a confidence corresponding to each prediction box; the image content in the edge box of the prediction box includes a target object recognized by the target detection model; and the confidence indicates a prediction probability that the target object is a power equipment. The target screening module is configured to perform fusion processing on the to-be-labeled infrared image and the to-be-labeled visible light image of the power site to obtain a to-be-labeled fusion image; perform background padding at an edge of the to-be-labeled infrared image to obtain a to-be-labeled infrared augmented image according to a size of the to-be-labeled infrared image and a size of the to-be-labeled visible light image, image content of the to-be-labeled visible light image, and image content of the to-be-labeled infrared image; the size of the to-be-labeled infrared augmented image is the same as the size of the to-be-labeled visible light image; a relative position of the image content of the to-be-labeled infrared image in the to-be-labeled infrared augmented image is the same as a relative position of the image content of the to-be-labeled infrared image in the to-be-labeled visible light image; and a target prediction box is determined in at least one prediction box for labeling a same target object in the plurality of single-modal labeled images according to the confidence. The fault diagnosis module is configured to determine a fault monitoring result of a corresponding power equipment according to the image content in the edge box of the target prediction box. The target screening module is specifically configured to: according to image content of the to-be-labeled infrared image, determine image regions corresponding to the image content of the to-be-labeled infrared image in a visible light labeled image and in a fusion labeled image as visible light labeled regions and fusion labeled regions respectively; determine all bounding boxes included in the visible light labeled regions as first bounding boxes, determine all bounding boxes included in the fusion labeled regions as second bounding boxes, and determine all bounding boxes included in the infrared augmented labeled image as third bounding boxes; for each third bounding box, determine, as candidate first bounding boxes, the first bounding boxes in which the similarity between the image content in the bounding box and the image content in the third bounding box is greater than or equal to a similarity threshold; for each third bounding box, determine, as first target bounding boxes, the bounding boxes with the maximum confidence in the third bounding box and the corresponding candidate first bounding boxes; determine, as candidate second bounding boxes, the second bounding boxes in which the similarity between the image content in the bounding box and the image content in the candidate first bounding box is greater than or equal to the similarity threshold, where the candidate first bounding box is a remaining bounding box in the visible light labeled region after the candidate first bounding boxes and the first target bounding boxes are removed; for each candidate first bounding box, determine, as second target bounding boxes, the bounding boxes with the maximum confidence in the candidate first bounding box and the candidate second bounding boxes; for each candidate third bounding box, determine, as third target bounding boxes, the second bounding boxes in which the similarity between the image content in the bounding box and the image content in each candidate third bounding box is less than the similarity threshold, where the candidate second bounding box is a remaining second bounding box in the fusion labeled region after the candidate second bounding boxes and the second target bounding boxes are removed, and the candidate third bounding box is a remaining third bounding box that is not determined as the first target bounding box; and determine the first target bounding boxes, the second target bounding boxes, and the third target bounding boxes as the target bounding boxes.
6. An electronic device, comprising: The electronic device includes a processor and a memory, the memory is used to store an application program, and the processor runs or executes a software program stored in the memory to enable the electronic device to implement the fault monitoring method of the power equipment according to any one of claims 1 to 4.
7. A computer readable storage medium characterized by, The computer readable storage medium is used to store program codes executed by the processor, and the program codes are used to implement the fault monitoring method of the power equipment according to any one of claims 1 to 4.
8. A computer program product, characterised in that, The computer program product contains computer instructions, when the computer instructions run on the electronic device, enable the electronic device to implement the fault monitoring method of the power equipment according to any one of claims 1 to 4.
Citation Information
Patent Citations
System equipment state monitoring method and related device
CN117994618A