State detection method and device of target equipment, electronic equipment and storage medium

Through multiple iterations training of the object detection model and the fault detection model, combined with adjustment of threshold values ​​and cluster analysis, the problem of low detection accuracy of the traditional dynamic object detection method in complex backgrounds and dynamic scenarios is solved, and high accuracy detection of the target equipment status during power facility inspection is achieved.

CN120235833APending Publication Date: 2025-07-01STATE GRID BEIJING ELECTRIC POWER CO +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510314577.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Traditional dynamic target detection methods have low detection accuracy in complex backgrounds and dynamic scenarios, making it difficult to effectively solve the problem of target equipment status detection accuracy in power facility inspection.

Method used

By acquiring the device image of the power system, the object detection model obtained by multiple iterative training determines the device position, and a clearer image is obtained based on the position, and the fault detection model is input to obtain the device status. This method improves the accuracy of dynamic object detection by adjusting threshold and clustering analysis.

Benefits of technology

It achieves the improvement of the accuracy of dynamic target detection, can more accurately identify and locate target equipment in the power system, and improves the accuracy and reliability of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235833A_ABST
    Figure CN120235833A_ABST
Patent Text Reader

Abstract

The invention discloses a state detection method and device of target equipment, electronic equipment and a storage medium. The method belongs to the field of power systems, and comprises the steps that a first equipment image of the power system is acquired, and the first equipment image comprises target equipment needing to be subjected to state detection; the first equipment image is input into a target detection model, the equipment position of the target equipment is obtained, the target detection model is a model obtained by conducting multiple iteration training on an initial detection model, adjustment thresholds used in different iteration training rounds are different, and the adjustment thresholds are used for determining whether the initial detection model is adjusted or not; based on the device position, obtaining a second device image of the target device, the number of noisy points of the second device image being smaller than the number of noisy points of the first device image; and inputting the second equipment image into the fault detection model to obtain the equipment state of the target equipment. According to the invention, the technical problem of low accuracy of dynamic target detection in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems, and in particular, to a method and device for detecting the state of a target device, an electronic device, and a storage medium. Background Art

[0002] In the power industry, with the rapid expansion of the power grid scale and the increase in the complexity of power facilities, the inspection and maintenance of power facilities have become an important link to ensure the safe and stable operation of the power grid. The traditional inspection method mainly relies on manual on-site inspection, which is not only inefficient but also has great safety hazards in complex or dangerous environments. In recent years, the combination of unmanned aerial vehicle (UAV) technology and machine vision has provided a new solution for power inspection. Through high-altitude inspection by UAVs and combined with video image analysis, remote, efficient, and comprehensive detection of power facilities can be achieved, and potential safety hazards can be detected in a timely manner.

[0003] Dynamic target detection is one of the important technologies in UAV inspection. However, traditional dynamic target detection methods, such as frame difference method, optical flow analysis method, and background difference method, although perform well in static scenes, have problems of low detection accuracy in complex backgrounds and dynamic scenes, such as power facility inspection.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for detecting the state of a target device, an electronic device, and a storage medium, so as to at least solve the technical problem of low accuracy of dynamic target detection in related technologies.

[0006] According to one aspect of the embodiments of the present invention, a method for detecting the state of a target device is provided, including: obtaining a first device image of a power system, where the first device image includes a target device that needs to be detected for its state; inputting the first device image into a target detection model to obtain the device position of the target device, where the target detection model is a model obtained by performing multiple iterative trainings on an initial detection model, and adjustment thresholds used in different iterative training rounds are different, and the adjustment threshold is used to determine whether to adjust the initial detection model; based on the device position, obtaining a second device image of the target device, where the number of noise points in the second device image is less than that in the first device image; and inputting the second device image into a fault detection model to obtain the device state of the target device.

[0007] Further, the method further includes: obtaining a sample image of a power system and an annotation result corresponding to the sample image, where the sample image includes sample devices, and the annotation result is used to represent the result obtained by annotating the sample devices; performing multiple iterative trainings on the initial detection model based on the sample image and the annotation result to obtain a target detection model, where the current adjustment threshold used in the first iterative training round of the multiple iterative trainings is a preset threshold, and the current adjustment threshold used in other iterative training rounds of the multiple iterative trainings is obtained by adjusting the historical adjustment threshold used in the previous iterative training round based on the current loss value determined in the iterative training round and the historical loss value determined in the previous iterative training round, and the other iterative training rounds are any iterative training rounds other than the first iterative training round in the multiple iterative trainings.

[0008] Further, performing multiple iterative trainings on the initial detection model based on the sample image and the annotation result to obtain a target detection model includes: in any iterative training round, inputting the sample image into the initial detection model or the detection model corresponding to the previous iterative training round to obtain an initial detection result; determining the current loss value based on the annotation result and the initial detection result; adjusting the initial detection model based on the current loss value and the current adjustment threshold to obtain the detection model corresponding to any iterative training round, where the detection model corresponding to the last iterative training round in the multiple iterative trainings is the target detection model.

[0009] Further, the method further includes: in response to the current loss value being less than or equal to the historical loss value, performing an increasing operation on the historical adjustment threshold to obtain the current adjustment threshold; in response to the current loss value being greater than the historical loss value, performing a decreasing operation on the historical adjustment threshold to obtain the current adjustment threshold.

[0010] Further, inputting the sample image into the initial detection model or the detection model corresponding to the previous iterative training round to obtain an initial detection result includes: using the initial detection model or the detection model corresponding to the previous iterative training round to detect the sample image to obtain multiple initial prior positions; determining the initial distances corresponding to different initial prior positions based on the multiple initial prior positions and the annotation result; clustering the multiple initial prior positions according to the initial distances to obtain a clustering result; and determining the initial detection result based on the clustering result.

[0011] Further, the initial detection model includes: a feature extraction unit, a feature fusion unit, and an initial detection unit. The sample image is detected using the initial detection model or the detection model corresponding to the previous iteration training round to obtain a plurality of initial prior positions, including: using the feature extraction unit to extract features from the sample image to obtain sample features, where different channels of the sample image correspond to different convolutional kernels of the feature extraction unit; using the feature fusion unit to perform feature fusion on the sample features to obtain fused features; inputting the fused features into the initial detection unit, and using the initial detection unit to detect the sample image to obtain a plurality of initial prior positions, where the weight assigned to the initial detection unit for the first preset type of target device is greater than the weight of the second preset type of target device.

[0012] Further, the method further includes: using a preset recognition model to recognize the sample device in the sample image to obtain an initial device recognition result; determining an identification loss value corresponding to the initial device recognition result based on the initial device recognition result and the annotation result; and determining whether the type of the sample device is the first preset type based on the identification loss value.

[0013] According to another aspect of the embodiments of the present invention, there is also provided a state detection device for a target device, including: a first acquisition module, configured to acquire a first device image of a power system, where the first device image includes a target device that needs to be subjected to state detection; a position determination module, configured to input the first device image into a target detection model to obtain the device position of the target device, where the target detection model is a model obtained by performing multiple iterative trainings on the initial detection model, and the adjustment thresholds used in different iterative training rounds are different, and the adjustment threshold is used to determine whether to adjust the initial detection model; a second acquisition module, configured to acquire a second device image of the target device based on the device position, where the number of noise points in the second device image is less than the number of noise points in the first device image; and a state determination module, configured to input the second device image into a fault detection model to obtain the device state of the target device.

[0014] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a memory storing an executable program; and a processor configured to run the program, where when the program runs, it executes the methods in the various embodiments of the present invention.

[0015] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present invention.

[0016] On the other hand, according to an embodiment of the present invention, there is also provided a computer program product, including a computer program which, when executed by a processor, implements the methods in various embodiments of the present invention.

[0017] On the other hand, according to an embodiment of the present invention, there is also provided a computer program product, including a non-volatile computer-readable storage medium storing a computer program which, when executed by a processor, implements the methods in various embodiments of the present invention.

[0018] On the other hand, according to an embodiment of the present invention, there is also provided a computer program which, when executed by a processor, implements the methods in various embodiments of the present invention.

[0019] In the embodiment of the present invention, the method includes obtaining a first device image of a power system; inputting the first device image into a target detection model to obtain the device position of a target device; based on the device position, obtaining a second device image of the target device; and inputting the second device image into a fault detection model to obtain the device state of the target device. By iteratively training an initial detection model multiple times, a target detection model with high dynamic target detection accuracy is obtained. The target detection model is used to identify the first device image, accurately output the device position of the target device, and then obtain a clearer and more accurate second device image based on the position of the target device. Finally, by inputting the second device image into the fault detection model, the device state of the target device with high accuracy is obtained, achieving the purpose of improving the detection ability and accuracy for targets of different scales, thus realizing the technical effect of improving the accuracy of dynamic target detection, and further solving the technical problem of low accuracy of dynamic target detection in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0021] Figure 1 is a flowchart of a method for detecting the state of a target device according to an embodiment of the present invention;

[0022] Figure 2 is a schematic diagram of the network structure of an optional improved YOLOv7 model according to an embodiment of the present invention;

[0023] Figure 3 is a schematic logical diagram of an optional method for detecting the state of a target device according to an embodiment of the present invention;

[0024] Figure 4Schematic diagram of a state detection device for a target device according to an embodiment of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] According to an embodiment of the present invention, an embodiment of a method for detecting the state of a target device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.

[0028] Figure 1 Is a flowchart of a method for detecting the state of a target device according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0029] Step S102, obtain a first device image of the power system, where the first device image includes a target device that needs to be detected for its state.

[0030] The above-mentioned first device image may refer to an initial image of the power facility site collected by a drone or other monitoring devices. The above-mentioned first device image includes various devices in the power system, such as transmission towers, wires, insulators, transformers, etc., as well as environmental information around the devices. The above-mentioned first device image is the original data source for state detection, anomaly recognition, and target detection, and subsequent image processing, feature extraction, model prediction, etc. operations will be based on this image.

[0031] The above target devices may refer to power devices that need to be concerned, detected, or analyzed in the above first device image. These devices may become the focus of the detection task due to their location, status, function, or potential abnormal conditions. For example, in the inspection task of power facilities, the above target devices may include at least one or more of the following: transmission towers, wires, insulators, transformers, etc., but are not limited thereto.

[0032] In an optional embodiment, the target device detection system (hereinafter referred to as the detection system) may use a drone as a data collection tool to collect multi-directional images or videos of the target devices in the power system and return them to the detection system, enabling the detection system to obtain the above first device image in real time. The flexibility and high-altitude perspective of the above drone ensure the comprehensiveness and high quality of the first device image, providing a solid data foundation for subsequent detections.

[0033] Step S104: Input the first device image into the target detection model to obtain the device positions of the target devices. The target detection model is a model obtained by iteratively training the initial detection model multiple times. Different adjustment thresholds are used in different iterative training rounds, and the adjustment threshold is used to determine whether to adjust the initial detection model.

[0034] The above target detection model may refer to the model after training the above initial detection model. The main task of the above target detection model is to locate and identify objects or devices in the power system in the image. Compared with the above initial detection model, the above target detection model integrates more functions, such as dynamic threshold adjustment, separable convolution, the introduction of Focal Loss, and model pruning techniques, but is not limited thereto. Through multiple iterative trainings of the initial detection model, the above target detection model can more accurately identify the target devices in the power system and determine their positions.

[0035] The above device positions may be the precise position information of the power devices located in the first device image by the above target detection model. This information is usually represented in the form of a bounding box. The bounding box is composed of the coordinates of the upper left corner and the lower right corner or the center point coordinates and the width and height parameters, which clearly define the specific position and size of the device in the image. In the power inspection scenario, identifying the device positions helps to further analyze the device status, such as whether there are abnormal conditions such as cracks and rust on the current device, but is not limited thereto.

[0036] The above-mentioned initial detection model can be the original state of the object detection model at the beginning of training. For example, the above-mentioned initial detection model can be based on a pre-trained YOLOv7 (You Only Look Once version 7, an object detection framework) model or other basic model architectures, but not limited to this. The parameters and structure of the above-mentioned initial detection model are usually relatively general and need to be trained through a dedicated power equipment detection task. Through multiple iterative trainings, the above-mentioned initial detection model can gradually adjust its parameters, learn the characteristics of power equipment, and finally be transformed into an object detection model that can efficiently and accurately perform the dynamic detection task of power equipment.

[0037] The above-mentioned adjustment threshold can be a parameter used to adjust the behavior of the object detection model during the iterative training process. The above-mentioned adjustment threshold can be used to determine whether the initial detection model or the model in the current iteration round should be adjusted based on the current training state.

[0038] In an alternative embodiment, the detection system can use the YOLOv7 model as the above-mentioned initial detection model and perform multiple iterative trainings on the above-mentioned initial detection model to obtain an object detection model. In the initial stage of iterative training, the performance of the model in the current iterative training round or the initial training model may be poor. At this time, the detection system needs to further train the detection model or the initial detection model in the current iterative training round according to the above-mentioned adjustment threshold. In the later stage of iterative training, if the performance of the detection model in the current iterative training round has met the predetermined target, at this time, the detection system can end the above-mentioned iterative training process based on the current adjustment threshold, thereby obtaining an object detection model, and further improving the accuracy of dynamic object detection.

[0039] Step S106, based on the device location, obtain a second device image of the target device, where the number of noise points in the second device image is less than that in the first device image.

[0040] The above-mentioned second device image can be a device image obtained based on the above-mentioned device location. The above-mentioned second device image reduces noise and focuses on the target device area, which is beneficial to improving the accuracy of dynamic object detection. The above-mentioned noise points can refer to any non-expected and irregular pixel value changes in the image. The above-mentioned noise points may reduce the image quality, affect the visual effect, and interfere with the performance of the image analysis algorithm.

[0041] In an alternative embodiment, based on the first device image in the foregoing steps, the detection system can obtain the exact position of the target device. Based on this position information, the drone can perform directional secondary acquisition, that is, specifically perform more detailed imaging of the target device and its surrounding area to generate a second device image. In this embodiment, this directional acquisition strategy can ensure the image quality of the second device image, reduce the interference of irrelevant backgrounds, improve the clarity of device details, and further reduce the impact of noise, thereby providing a data basis for more accurately implementing dynamic target detection.

[0042] Step S108: Input the second device image into the fault detection model to obtain the device state of the target device.

[0043] The above-mentioned fault detection model can be a model for identifying and analyzing potential faults or abnormal states of power equipment. The above-mentioned device state can refer to the working condition or health condition information of the power equipment obtained through the analysis of the fault detection model. For example, the above-mentioned device state can include at least one or more of the following: whether the device is operating normally, whether there are cracks, specific faults and abnormal conditions of the power equipment, etc., but not limited thereto. The accurate identification of the above-mentioned device state is important for the preventive maintenance and rapid fault response of the power system. The above-mentioned device state can help the operation and maintenance personnel timely discover and solve potential problems, avoid power failures or safety accidents, and thus improve the reliability and efficiency of the power system operation.

[0044] In an alternative embodiment, the number of noise points in the above-mentioned second device image is lower than that of the first device image, enabling the fault detection model to more accurately identify and extract features related to the device state, such as the width of cracks, the degree of corrosion, the displacement of components, etc. Since the noise is effectively suppressed, the extraction of these features is more accurate, reducing false alarms and missed detections, and improving the accuracy and reliability of fault detection. By inputting the second device image into the above-mentioned fault detection model, it is possible to quickly locate and detect the faults of the target device. Compared with traditional fault detection methods, the method in this embodiment can instantaneously and accurately feedback the device state, that is, shorten the fault detection cycle and improve the accuracy of dynamic target detection.

[0045] In an embodiment of the present invention, the following method is adopted: obtaining a first device image of a power system; inputting the first device image into a target detection model to obtain the device position of a target device; based on the device position, obtaining a second device image of the target device; and inputting the second device image into a fault detection model to obtain the device state of the target device. By iteratively training an initial detection model multiple times to obtain a target detection model with high dynamic target detection accuracy, and using the above target detection model to identify the first device image, accurately outputting the device position of the target device, and then obtaining a clearer and more accurate second device image based on the position of the target device. Finally, by inputting the second device image into the fault detection model, the device state of the target device with high accuracy is obtained, achieving the purpose of improving the detection ability and accuracy of targets of different scales, thus realizing the technical effect of improving the accuracy of dynamic target detection, and further solving the technical problem of low accuracy of dynamic target detection in the related art.

[0046] Further, the method further includes: obtaining a sample image of the power system and the corresponding annotation result, where the sample image includes a sample device, and the annotation result is used to represent the result obtained by annotating the sample device; based on the sample image and the annotation result, performing multiple iterative trainings on the initial detection model to obtain a target detection model, where the current adjustment threshold used in the first iteration training round of the multiple iterative trainings is a preset threshold, and the current adjustment threshold used in other iterative training rounds of the multiple iterative trainings is obtained by adjusting the historical adjustment threshold used in the previous iteration training round based on the current loss value determined in the iterative training round and the historical loss value determined in the previous iteration training round, and the other iterative training rounds are any one of the multiple iterative training rounds except the first iteration training round.

[0047] The above sample image can be an image set collected from the power system for training and validating the detection model. These images contain different states and environmental conditions of various devices in the power system.

[0048] The above annotation result can be detailed identification and classification information for the power devices included in the sample image. The above annotation result can generate a bounding box and a class label for each power device or abnormal state, helping the detection model or the initial detection model in the current iterative training round to understand the specific position and state of the device in the image. The annotation result is the guiding information for model training, enabling the model to learn the features and fault patterns of the device.

[0049] The above sample devices may refer to the devices in the power system that appear in the sample images and need to be detected and classified. These devices may be in a normal state or a fault state, and their images and annotation information constitute the training set to help the detection model or the initial detection model in the current iterative training round learn how to identify and classify different devices and states.

[0050] The above preset threshold may be set at the initial stage of model training and is used as a parameter for training the above initial training model.

[0051] The above current loss value may be the value of the loss function calculated by the current detection model during the training stage in the current iterative training round. The loss function is used to evaluate the difference between the prediction result of the current detection model and the actual annotation result. The smaller the current loss value, the closer the prediction result of the current detection model is to the real data.

[0052] The above historical loss value may be the value of the loss function calculated in the previous iterative training round, and it is important reference information in the model training process. Based on the historical loss value and the current loss value, the detection model in the current training round can dynamically adjust the parameters of the model to improve the training effect.

[0053] In an alternative embodiment, the detection system can collect images through devices such as drones to obtain sample images containing different backgrounds, lighting conditions, and device states. Subsequently, the staff can carefully annotate these sample images, including information such as device type, location, and state, to generate accurate annotation results. This annotation process provides crucial training data for model training, ensuring that the initial detection model can learn the typical features and potential fault patterns of power equipment during the training process. In the first iteration training round, the detection system can use a preset CIoU Loss (Complete Intersection over Union Loss) threshold as the current adjustment threshold to train the initial detection model. The setting of the above preset threshold needs to be based on preliminary experiments or expert experience, aiming to ensure that the model can explore and learn a wide range of features of power equipment at the initial stage of training. The training in this round mainly enables the model to have a preliminary understanding of the basic form and location of power equipment, laying a foundation for subsequent optimization. Starting from the second iteration training round, the training of the detection model in the current round will adopt a dynamically adjusted CIoU Loss threshold mechanism. Specifically, the current adjustment threshold will be adjusted based on the current loss value determined in the current iteration training and the historical loss value determined in the previous iteration training. In this embodiment, the above mechanism can automatically adjust the looseness or strictness of the threshold according to the training state of the model and the change of the loss function, thereby improving the training effect of the detection model. After each iteration training, the detection system can evaluate the performance of the detection model corresponding to the current iteration process, such as detection accuracy, computational efficiency, etc. If the performance of the detection model corresponding to the current iteration process does not meet the expected standard, the iteration training will continue. If the performance of the detection model corresponding to the current iteration process has met the expected standard, the detection system will end the iteration training to obtain the target detection model.

[0054] Furthermore, multiple iterative trainings are performed on the initial detection model based on the sample images and the annotation results to obtain the target detection model, including: in any iteration training round, input the sample images into the initial detection model or the detection model corresponding to the previous iteration training round to obtain the initial detection results; determine the current loss value based on the annotation results and the initial detection results; adjust the initial detection model based on the current loss value and the current adjustment threshold to obtain the detection model corresponding to any iteration training round, where the detection model corresponding to the last iteration training round in the multiple iterative trainings is the target detection model.

[0055] In an alternative embodiment, the detection system may first input a sample image of the power system into the initial detection model or the detection model corresponding to the previous iteration training round. This process utilizes the forward propagation ability of the model to generate an initial detection result for the sample device. The above initial detection result may include the positioning bounding box of the device and possible class labels. Since the detection model in iterative training can adjust its parameters after each iterative training, starting from the second iteration, the corresponding detection model used in each round of iterative training is a version adjusted based on the results of the previous iterative training. This ensures that during multiple iterative training processes, the recognition accuracy of the trained detection model can be continuously improved. Subsequently, the detection system can calculate the current loss value based on the annotation result and the initial detection result of the sample image. The above loss value can not only reflect the intersection over union (IoU) of the bounding box and the prediction box, but also reflect the consistency of the center point distance and the aspect ratio. Finally, based on the current loss value and the current adjustment threshold, the detection system can adjust the parameters of the detection model corresponding to the previous iteration round, thereby improving the detection accuracy of the detection model corresponding to the current iteration round.

[0056] Further, the method further includes: in response to the current loss value being less than or equal to the historical loss value, performing an increasing operation on the historical adjustment threshold to obtain the current adjustment threshold; in response to the current loss value being greater than the historical loss value, performing a decreasing operation on the historical adjustment threshold to obtain the current adjustment threshold.

[0057] The above historical adjustment threshold may refer to the adjustment threshold used in the previous round of iterative training. The above historical adjustment threshold is dynamically calculated based on the statistical results of the previous round of training. For example, the above statistical results may be the average CIoU Loss value and the standard deviation of the CIoU Loss value, etc., but are not limited thereto.

[0058] The above current adjustment threshold may refer to the adjustment threshold used in the current round of iterative training. The above current adjustment threshold can be adjusted according to the comparison result between the current loss value calculated in this round of training and the historical loss value of the previous round of training.

[0059] In an alternative embodiment, the detection system can calculate the current loss value in the current iteration training round, that is, the degree of difference between the prediction result and the annotation result of the corresponding detection model in the current iteration training round. The above-mentioned current loss value can include various factors such as bounding box localization deviation and class recognition error, and is an important indicator for adjusting the corresponding detection model in the current iteration training round. If the current loss value is less than or equal to the historical loss value, this usually indicates that the detection model corresponding to the current iteration training round is converging stably or has reached a relatively ideal detection state. At this time, the detection system can obtain the current adjustment threshold by increasing the historical adjustment threshold, so as to encourage the model to further explore and improve, and allow a certain degree of detection error during the exploration and improvement process, so that the detection model corresponding to the subsequent iteration training round can learn more extensive target features and try to solve some boundary problems or small targets that are difficult to detect, thereby improving the overall detection performance of the finally generated target detection model. If the current loss value is greater than the historical loss value, this indicates that the detection model corresponding to the current iteration training round may encounter a training bottleneck or start to show signs of overfitting. At this time, the detection system can perform a reduction operation on the historical adjustment threshold to obtain a more stringent current adjustment threshold. The above steps can make the detection model corresponding to the subsequent iteration training round more focused on reducing the current detection error, avoid excessive exploration of complex but possibly irrelevant target features, thereby correcting the training direction and preventing overfitting, and then ensuring the detection accuracy of the finally generated target detection model.

[0060] For example, the calculation method of the loss threshold of the detection model corresponding to any iteration training round can be shown as the following formula:

[0061] τ(t) = μ + k·σ.

[0062] In the formula, τ(t) is the threshold of CIoU Loss in the t-th round of training, μ is the average value of all CIoU Loss values, σ is the standard deviation of CIoU Loss in all rounds, and k is a hyperparameter that controls the degree of deviation of the threshold from the average value. When the average CIoU Loss value in the t-th round of training is less than or equal to the average CIoU Loss value in the (t - 1)-th round of training, it means that the detection model corresponding to the current round is learning. At this time, the hyperparameter k can be appropriately relaxed to encourage the detection model corresponding to the current round to explore more possibilities. When the average CIoU Loss value in the t-th round of training is greater than the average CIoU Loss value in the (t - 1)-th round of training, it means that the model may encounter difficulties. At this time, the hyperparameter k can be tightened to make the model focus on reducing existing errors and avoid prematurely trying more complex bounding box predictions.

[0063] Further, input the sample image into the initial detection model or the detection model corresponding to the previous iteration training round to obtain the initial detection result, including: using the initial detection model or the detection model corresponding to the previous iteration training round to detect the sample image to obtain a plurality of initial prior positions; determining the initial distances corresponding to different initial prior positions based on the plurality of initial prior positions and the annotation result; clustering the plurality of initial prior positions according to the initial distances to obtain a clustering result; and determining the initial detection result based on the clustering result.

[0064] The above initial prior positions may be the possible positions of the target preset in the initial detection model or the detection model corresponding to the previous iteration training round. The above initial distance may be an index for measuring the difference between the initial prior position predicted by the model and the true target position in the above annotation result.

[0065] In an alternative embodiment, in each round of iterative training, the detection model corresponding to the current iteration round can detect the input sample image to generate a series of initial prior positions, that is, the bounding boxes predicted by the detection model corresponding to the current iteration round. These bounding boxes reflect the preliminary estimation of the possible positions of the power equipment in the image by the detection model corresponding to the current iteration round and are the basis for subsequent calculation of the initial distance and clustering. For each initial prior position, the detection system can calculate the initial distance between the initial prior position and the true position in the annotation result. The above initial distance calculation not only includes the position deviation between the predicted bounding box and the true box, but may also include the differences in size and aspect ratio between the predicted bounding box and the true box, thereby providing a quantitative index for evaluating the prediction accuracy of the bounding box for the detection model corresponding to the current iteration round. Then, the detection system can perform clustering analysis on the plurality of initial prior positions according to the calculated initial distances to identify the prior box categories that can better cover the true size distribution of the power equipment. The above clustering result can reflect the important features of the power equipment at different scales, thereby helping the detection model corresponding to the current iteration round to focus more on these important features in subsequent training, and further improving the training effect. Finally, after determining the clustering result, the detection model corresponding to the current iteration round can determine the initial detection result based on these optimized prior box distributions. This process utilizes the information provided by the clustering analysis to guide the detection model corresponding to the current iteration round to more accurately locate and identify the power equipment.

[0066] For example, the detection system can calculate the above initial distance according to the following formula:

[0067] d(box,centroid)=1 - IOU(box,centroid).

[0068] Wherein, d(box, centroid) represents the above-mentioned initial distance, IOU represents the intersection over union calculation, box represents the prior box, and centroid represents the target center. In addition, the detection system can also calculate the clustering target according to the following formula:

[0069] K aim = min ∑[1 - IOU(box, truth)].

[0070] Wherein, K aim represents the clustering target, box represents the prior box, truth represents the ground truth box, ∑[1 - IOU(box, truth)] represents the summation of the matching degree between the prior box and the ground truth box, and min represents taking the smaller value of the above summation result. The meanings of other symbols in the formula are the same as those in the foregoing formula and will not be elaborated here.

[0071] Furthermore, the initial detection model includes: a feature extraction unit, a feature fusion unit, and an initial detection unit. The sample image is detected using the initial detection model or the detection model corresponding to the previous iteration training round to obtain multiple initial prior positions, including: using the feature extraction unit to extract features from the sample image to obtain sample features, wherein different channels of the sample image correspond to different convolutional kernels of the feature extraction unit; using the feature fusion unit to perform feature fusion on the sample features to obtain fused features; inputting the fused features into the initial detection unit, and using the initial detection unit to detect the sample image to obtain multiple initial prior positions, wherein the weight assigned to the initial detection unit for the first preset type of target device is greater than the weight of the second preset type of target device.

[0072] The above-mentioned first preset type may be the type of target device that is not easily detected and recognized in power inspection. The above-mentioned second preset type may be the type of target device that is more easily detected and recognized in power inspection.

[0073] In an alternative embodiment, the detection system first processes the above sample image using a feature extraction unit, which may include multiple convolutional kernels, each corresponding to a different channel of the sample image. For example, the detection system may employ separable convolution, i.e., a combination of depthwise convolution and pointwise convolution, to reduce computational complexity and the number of model parameters. This lightweight design not only improves the running efficiency of the model on resource-constrained platforms such as drones but also ensures that the model can quickly extract features from images of power facilities to generate sample features. After obtaining the sample features, the detection system can comprehensively process these features using a feature fusion unit to achieve the fusion of features at different levels. For example, the detection system may adopt feature fusion techniques such as the CBS (Convolution, Batch Normalization, and SiLU, convolution, batch normalization, and SiLU activation function) module and the ELAN (Enhanced Lightweight Aggregation Network) module to effectively enhance feature representation and improve the model's understanding ability of complex power environments. In addition, the detection system may also introduce the SPPCSPC (Spatial Pyramid Pooling Concatenated with Cross Stage Partial) module to increase the receptive field of the initial detection model, thereby helping the initial detection model identify subtle changes and detailed features of power equipment in the image, which is crucial for detecting detailed anomalies such as cracks and rust on the detection equipment. Through feature fusion, the initial detection model can combine feature information at different scales and depths to generate more comprehensive and accurate fused features, providing strong support for subsequent detection tasks. The fused features are then input into the initial detection unit to generate multiple initial prior positions. In this embodiment, different weights are also set for different types of target devices. Specifically, target devices that are difficult to identify have higher weights, while devices that are easier to identify have lower weights. This weight assignment strategy can reduce the weights of easily classified samples during the iterative training process, enabling the detection model in training to concentrate more resources on processing difficult-to-classify sample devices, thereby improving the model's detection performance for small targets.

[0074] For example, the detection system may introduce Focal Loss into the original classification loss function of the YOLOv7 model. At this time, the loss function of the initial detection model can be shown as follows:

[0075]

[0076] In the formula, L Focal-CE represents the loss function, C is the total number of categories in the classification task, i is the index, and αt To balance the weight factor of positive and negative samples, p t is the probability that the model predicts the correct class, γ is a hyperparameter in Focal Loss, used to control the decreasing speed of the loss weight of easy-to-classify samples, y i is the true class, and p i is the probability of each predicted class.

[0077] For ease of understanding, Figure 2 is a schematic diagram of the network structure of an optional improved YOLOv7 model according to an embodiment of the present invention, as Figure 2As shown, the structure includes a backbone network, an input, a feature fusion layer, and a detection layer. Among them, the input is the input end of the model, which receives the processed image data. The backbone network is used to extract features from the input image. The backbone network may include ELAN1, ELAN2, ELAN3, and ELAN4 modules, MP-11, MP-12, and MP-13 modules, and 4×Depthwise conv (Depthwise Convolution) modules. Among them, the ELAN1, ELAN2, ELAN3, and ELAN4 modules are used to enhance the feature extraction ability while keeping the calculation lightweight. The MP-11, MP-12, and MP-13 modules are used to reduce the spatial dimension of the feature map, thereby reducing the amount of calculation and the number of parameters while maintaining the robustness of feature detection. The 4×Depthwise conv is used to reduce the amount of calculation and the number of parameters while maintaining the expressiveness of the model. The 4×Depthwise conv module receives the input data and further transmits it to the ELAN1 module. The ELAN1 module processes the data and then transmits the data to the MP-11 module. Subsequently, the MP-11 module passes the processed data to the ELAN2 module. At this time, the ELAN2 module processes the received data and passes the processed data to the MP-12 module and the Depthwise conv1 module of the feature fusion layer. After the MP-12 module processes the received data, it further transmits it to the ELAN3 module. Subsequently, the ELAN3 module passes the processed data to the MP-13 module and the Depthwise conv2 module of the feature fusion layer. After receiving the data, the MP-13 module processes the data and transmits it to the ELAN4 module. Subsequently, the ELAN4 module passes the processed data to the SPPCSPC module of the feature fusion layer. The feature fusion layer is used to fuse feature information at different levels to enhance the multi-scale object detection ability of the model. The feature fusion layer includes Depthwise conv1, Depthwise conv2 modules, SPPCSPC module, CPS (Channel Pruning Strategy) module, CBS module, UPSample1 (Upsampling) module, UPSample2 module, cat1 (Concatenation) module, cat2 module, ELAN5, ELAN6, ELAN7, ELAN8 modules, MP-21 and MP-22 modules. Among them, the Depthwise conv1 and Depthwise conv2 modules are used to reduce the amount of calculation and the number of parameters while maintaining the efficient processing of the input features. The SPPCSPC module is used to enhance the receptive field of the model and capture features at different scales. The CPS module is used to extract and fuse multi-scale features by constructing a pyramid structure.The CBS module is used for feature extraction, normalization, and activation. The UPSample1 module and the UPSample2 module are used to upsample the low-resolution feature maps to a higher resolution for easy fusion with features at higher scales. The cat1 module and the cat2 module are used to concatenate the feature maps from different modules to merge multi-scale and multi-level information. The functions of the ELAN5, ELAN6, ELAN7, and ELAN8 modules are similar to their functions in the backbone network. The MP-21 and MP-22 modules are used to reduce the computational amount and the number of parameters while maintaining the robustness of the model for object detection. The Depthwise conv1 module receives the data transmitted by the ELAN2 module in the backbone network and further transmits it to the cat1 module after processing the data. Similarly, the Depthwise conv2 module receives the data transmitted by the ELAN3 module in the backbone network and further transmits it to the cat2 module. The SPPCSPC module is used to receive the data passed by the ELAN4 module in the backbone network and further pass it to the CPS module and the MP-22 module in the feature fusion layer. After receiving the data, the CPS module passes it to the UPSample2 module, and the UPSample2 module then passes the data to the cat2 module. The cat2 module processes the data transmitted by the Depthwise conv2 module and the UPSample2 module and outputs it to the ELAN6 module in the feature fusion layer. The ELAN6 module then passes the processed data to the CBS module and the MP-21 module in the detection layer. After receiving the data, the CBS module passes it to the UPSample1 module, and then the UPSample1 passes the data to the cat1 module. After processing the data from the Depthwise conv1 module and the UPSample1, the cat1 module passes it to the ELAN6 module. After processing, the ELAN6 module passes it to the MP-21 module and the REP1 (Repetitive Element Pruning) module in the detection layer. The MP-21 module processes the data from the ELAN6 module and the ELAN5 module and sends it to the ELAN7 module. After processing the data, the ELAN7 module sends it to the MP-22 module and the REP2 module in the detection layer. The MP-22 module processes the data from the SPPCSPC module and the ELAN7 module and passes it to the ELAN8 module. After processing, the ELAN8 module sends the data to the REP3 module in the detection layer. The detection layer is used to generate the bounding boxes and class predictions of the targets. The detection layer includes the REP1, REP2, and REP3 modules. The functions of the above REP1, REP2, and REP3 modules are to deeply process the input feature maps by repeatedly performing the same detection operations to improve the accuracy and stability of object detection.Among them, the REP1 module receives data from ELAN6, processes it, and then performs detection and output. Similarly, the REP2 module receives data from ELAN7, processes it, and then performs detection and output. The REP3 module receives data from ELAN8, processes it, and then performs detection and output.

[0078] Furthermore, the method further includes: using a preset recognition model to recognize the sample device in the sample image to obtain an initial device recognition result; determining an identification loss value corresponding to the initial device recognition result based on the initial device recognition result and the annotation result; and determining whether the type of the sample device is a first preset type based on the identification loss value.

[0079] The above-mentioned preset recognition model can be a pre-set model for recognizing target device categories. The above-mentioned initial device recognition result can refer to the device type in the sample device obtained by the above-mentioned preset recognition model after recognizing the sample image. The above-mentioned identification loss value can be a parameter value for determining whether the type of the sample device is a first preset type.

[0080] In an alternative embodiment, the sample image is first processed by a preset recognition model to generate a preliminary recognition result of the device type of the sample device in the image. Based on the comparison between the initial device recognition result generated by the model and the annotation result, the detection system can calculate the above-mentioned identification loss value, and further determine whether the type of the sample device is a first preset type based on the magnitude of the above-mentioned identification loss value. Specifically, if the above-mentioned identification loss value is large, it indicates that the sample device is not easily recognized. Therefore, the device type of the sample device should be classified as the above-mentioned first preset type.

[0081] For ease of understanding, Figure 3 is a logical schematic diagram of an alternative method for detecting the state of a target device according to an embodiment of the present invention. As Figure 3 shown, first, an inspection device is used to collect and generate a target image to be detected. Subsequently, the detection system can perform training and optimization of dynamic target detection on the inspection data based on the improved YOLOv7 model. Then, the detection system can reduce the model parameter quantity by designing prior boxes and threshold adjustment, introducing separable convolutions, and enhancing the training weight of small samples by using classification loss combined with focal loss, and at the same time perform structural pruning on the improved model. Finally, the detection system can deploy the trained dynamic detection model to perform task inspection of the power system.

[0082] According to an embodiment of the present invention, an embodiment of a device for detecting the state of a target device is provided. It should be noted that this device can be used to execute the above-mentioned method for detecting the state of a target device. The specific implementation manner and application scenario are the same as those of the above embodiment and will not be elaborated here. Figure 4Schematic diagram of a state detection device for a target device according to an embodiment of the present invention, as Figure 4 shown, the device includes:

[0083] A first acquisition module 402, configured to acquire a first device image of a power system, where the first device image includes a target device that needs to be subjected to state detection.

[0084] A position determination module 404, configured to input the first device image into a target detection model to obtain the device position of the target device, where the target detection model is a model obtained by performing multiple iterative trainings on an initial detection model, and adjustment thresholds used in different iterative training rounds are different, and the adjustment threshold is used to determine whether to adjust the initial detection model.

[0085] A second acquisition module 406, configured to acquire a second device image of the target device based on the device position, where the number of noise points in the second device image is less than the number of noise points in the first device image.

[0086] A state determination module 408, configured to input the second device image into a fault detection model to obtain the device state of the target device.

[0087] Further, the device further includes: a third acquisition module, configured to acquire a sample image of the power system and an annotation result corresponding to the sample image, where the sample image includes a sample device, and the annotation result is used to represent the result obtained by annotating the sample device; a model training module, configured to perform multiple iterative trainings on the initial detection model based on the sample image and the annotation result to obtain the target detection model, where the current adjustment threshold used in the first iterative training round in the multiple iterative trainings is a preset threshold, and the current adjustment threshold used in other iterative training rounds in the multiple iterative training rounds is obtained by adjusting the historical adjustment threshold used in the previous iterative training round based on the current loss value determined in the iterative training round and the historical loss value determined in the previous iterative training round, and the other iterative training rounds are any one of the multiple iterative training rounds except the first iterative training round.

[0088] Further, the model training module is further configured to: input the sample image into the initial detection model or the detection model corresponding to the previous iterative training round to obtain an initial detection result; determine the current loss value based on the annotation result and the initial detection result; adjust the initial detection model based on the current loss value and the current adjustment threshold to obtain the detection model corresponding to any one of the iterative training rounds, where the detection model corresponding to the last iterative training round in the multiple iterative trainings is the target detection model.

[0089] Further, the device further includes: a threshold increasing module, configured to perform an increasing operation on a historical adjustment threshold to obtain a current adjustment threshold in response to the current loss value being less than or equal to the historical loss value; a threshold decreasing module, configured to perform a decreasing operation on the historical adjustment threshold to obtain a current adjustment threshold in response to the current loss value being greater than the historical loss value.

[0090] Further, the model training module is further configured to: use the initial detection model or the detection model corresponding to the previous iteration training round to detect the sample image to obtain a plurality of initial prior positions; determine the initial distances corresponding to different initial prior positions based on the plurality of initial prior positions and the annotation results; cluster the plurality of initial prior positions according to the initial distances to obtain a clustering result; and determine an initial detection result based on the clustering result.

[0091] Further, the initial detection model includes: a feature extraction unit, a feature fusion unit, and an initial detection unit. The model training module is further configured to: use the feature extraction unit to extract features from the sample image to obtain sample features, where different channels of the sample image correspond to different convolutional kernels of the feature extraction unit; use the feature fusion unit to perform feature fusion on the sample features to obtain fused features; input the fused features into the initial detection unit, and use the initial detection unit to detect the sample image to obtain a plurality of initial prior positions, where the weight assigned to the target device of the first preset type by the initial detection unit is greater than the weight of the target device of the second preset type.

[0092] Further, the device further includes: a first recognition module, configured to use a preset recognition model to recognize the sample device in the sample image to obtain an initial device recognition result; a first determination module, configured to determine an identification loss value corresponding to the initial device recognition result based on the initial device recognition result and the annotation result; and a second determination module, configured to determine whether the type of the sample device is the first preset type based on the identification loss value.

[0093] An embodiment of the present application further provides an electronic device, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the methods in the various embodiments of the present invention.

[0094] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present invention.

[0095] An embodiment of the present application further provides a computer program product, including a computer program that implements the methods in the various embodiments of the present invention when executed by a processor.

[0096] Embodiments of the present application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program, where the computer program, when executed by a processor, implements the methods in various embodiments of the present invention.

[0097] Embodiments of the present application also provide a computer program, where the computer program, when executed by a processor, implements the methods in various embodiments of the above-mentioned present invention.

[0098] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0099] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.

[0100] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0101] In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0102] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0103] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for detecting the state of a target device, characterized in that: include: Acquire a first device image of the power system, wherein the first device image includes a target device that needs to be detected; Inputting the first device image into a target detection model to obtain a device position of the target device, wherein the target detection model is a model obtained by performing multiple iterative training on an initial detection model, and different adjustment thresholds are used in different iterative training rounds, and the adjustment threshold is used to determine whether to adjust the initial detection model; Based on the device position, acquiring a second device image of the target device, wherein the number of noise points in the second device image is less than the number of noise points in the first device image; The second device image is input into a fault detection model to obtain the device status of the target device.

2. The method according to claim 1, characterized in that The method further comprises: Acquire a sample image of the power system and a labeling result corresponding to the sample image, wherein the sample image includes a sample device, and the labeling result is used to characterize a result obtained by labeling the sample device; The initial detection model is trained for multiple times based on the sample image and the annotation result to obtain the target detection model, wherein the current adjustment threshold used in the first iterative training round of the multiple iterative training is a preset threshold, the current adjustment threshold used in other iterative training rounds of the multiple iterative training is obtained by adjusting the historical adjustment threshold used in the previous iterative training round based on the current loss value determined in the iterative training round and the historical loss value determined in the previous iterative training round, and the other iterative training rounds are any iterative training rounds in the multiple iterative training except the first iterative training round.

3. The method according to claim 2, characterized in that The initial detection model is iteratively trained multiple times based on the sample image and the annotation result to obtain the target detection model, including: In any iterative training round, the sample image is input into the initial detection model or the detection model corresponding to the previous iterative training round to obtain an initial detection result; Determine a current loss value based on the labeling result and the initial detection result; Based on the current loss value and the current adjustment threshold, the initial detection model is adjusted to obtain the detection model corresponding to any one of the iterative training rounds, wherein the detection model corresponding to the last iterative training round in the multiple iterative trainings is the target detection model.

4. The method according to claim 2 or 3, characterized in that: The method further comprises: In response to the current loss value being less than or equal to the historical loss value, increasing the historical adjustment threshold to obtain the current adjustment threshold; In response to the current loss value being greater than the historical loss value, the historical adjustment threshold is reduced to obtain the current adjustment threshold.

5. The method according to claim 3, characterized in that: Input the sample image into the initial detection model or the detection model corresponding to the last iterative training round to obtain an initial detection result, including: Detecting the sample image using the initial detection model or the detection model corresponding to the last iterative training round to obtain a plurality of initial prior positions; Based on the multiple initial a priori positions and the marking results, determining initial distances corresponding to different initial a priori positions; Clustering the multiple initial a priori positions according to the initial distance to obtain a clustering result; Based on the clustering result, the initial detection result is determined.

6. The method according to claim 5, characterized in that The initial detection model includes: a feature extraction unit, a feature fusion unit and an initial detection unit. The sample image is detected using the initial detection model or the detection model corresponding to the last iterative training round to obtain multiple initial prior positions, including: Using the feature extraction unit, extracting features from the sample image to obtain sample features, wherein different channels of the sample image correspond to different convolution kernels of the feature extraction unit; Using the feature fusion unit, the sample features are fused to obtain fused features; The fused features are input into the initial detection unit, and the sample image is detected using the initial detection unit to obtain the multiple initial prior positions, wherein the weight assigned by the initial detection unit to the target device of the first preset type is greater than the weight of the target device of the second preset type.

7. The method according to claim 6, characterized in that The method further comprises: Using a preset recognition model, the sample device in the sample image is identified to obtain an initial device recognition result; Determining an identification loss value corresponding to the initial device identification result based on the initial device identification result and the labeling result; Based on the recognition loss value, it is determined whether the type of the sample device is the first preset type.

8. A state detection device for a target device, characterized in that: include: A first acquisition module, used to acquire a first device image of the power system, wherein the first device image includes a target device that needs to be detected; a position determination module, configured to input the first device image into a target detection model to obtain a device position of the target device, wherein the target detection model is a model obtained by performing multiple iterative training on an initial detection model, and different adjustment thresholds are used in different iterative training rounds, and the adjustment threshold is used to determine whether to adjust the initial detection model; A second acquisition module, configured to acquire a second device image of the target device based on the device position, wherein the number of noise points in the second device image is smaller than the number of noise points in the first device image; The state determination module is used to input the second device image into the fault detection model to obtain the device state of the target device.

9. An electronic device, characterized in that: include: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 6 when running.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 6.