Unmanned aerial vehicle inspection method and device for on-line detection of power system, and medium

By conducting YOLOv8 model training and detection of power facility images, the problems of low efficiency and safety risks of power system inspection are solved, real-time defect identification and rapid response are achieved.

CN120431019APending Publication Date: 2025-08-05GUANGXI POWER GRID CO LTD FANGCHENGGANG POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510429305.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing power system inspection methods rely on manual inspections with low efficiency, high cost and security risks. The existing drone inspection methods have data transmission delays and privacy leakage problems, and real-time fault warning cannot be achieved.

Method used

The YOLOv8 defect detection model is used to train and detect the power facility images captured by the drone, and the model is trained by acquiring historical image sets, and defect recognition is used to identify the defects using feature extraction and fusion modules, and defect information is determined in real-time image data.

Benefits of technology

It realizes rapid response and accurate identification of the power system, improves inspection efficiency and accuracy, and reduces the cost and risks of manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431019A_ABST
    Figure CN120431019A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power systems, and provides a power system online detection-oriented unmanned aerial vehicle inspection method and device and a medium, and the method comprises the steps: firstly, obtaining a historical image set of an unmanned aerial vehicle in an inspection target power region; secondly, determining a training set based on the historical image set; and training a to-be-trained YOLOv8 defect detection model based on the training set to obtain a trained YOLOv8 defect detection model, and then, obtaining real-time image data of the unmanned aerial vehicle in the inspection target power region, and finally, determining defect information about the target power region based on the trained YOLOv8 defect detection model and the real-time image data. According to the embodiment of the invention, efficient defect detection is carried out by using the trained YOLOv8 model, rapid response and accurate identification of online detection of the power system are realized, the inspection efficiency and accuracy are effectively improved, and the cost and risk of manual inspection are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of power systems. Specifically, it relates to an unmanned aerial vehicle (UAV) inspection method, device, and medium for online detection of power systems. Background Art

[0002] With the continuous expansion of the scale and complexity of power systems, the number and types of power equipment are increasing day by day. How to ensure the safety, reliability, and efficient operation of power systems has become a major challenge faced by modern power engineering. Faults in power systems may not only cause power outages, affecting residential life and industrial production, but may also result in economic losses and even casualties in severe cases. Therefore, to ensure the stable operation of power system equipment, timely and effective inspection and fault diagnosis are required, especially regular inspections in high-risk areas and inaccessible places.

[0003] Currently, the inspection of power systems mainly relies on manual inspection and regular maintenance. Although these methods can ensure a certain inspection frequency, there are also some obvious drawbacks: First, manual inspection usually requires a large amount of time and manpower, with low work efficiency, and is prone to subjective judgment errors due to manual operations, making it difficult to ensure the stability of inspection quality. Second, power equipment is often widely distributed, especially in high-risk or inaccessible areas such as high-voltage lines and transmission towers. Manual inspection not only increases safety risks but also raises costs and difficulties. At the same time, the defect detection methods in related technologies mostly rely on uploading image data to the cloud for analysis and processing. This method has problems such as high data transmission latency and large network bandwidth requirements, and cannot respond to fault warnings in real time; moreover, there may be privacy leakage and security issues during the data transmission process. Summary of the Invention

[0004] The embodiments of the present disclosure at least provide an unmanned aerial vehicle (UAV) inspection method, device, and medium for online detection of power systems. By using the trained YOLOv8 model for efficient defect detection, it realizes the rapid response and accurate identification of online detection of power systems, effectively improves the efficiency and accuracy of inspection, and reduces the costs and risks of manual inspection.

[0005] The embodiments of the present disclosure provide an unmanned aerial vehicle (UAV) inspection method for online detection of power systems, including:

[0006] Obtain a historical image set of the UAV in the target power area to be inspected; wherein, the historical image set includes multiple labeled historical images of the target power area;

[0007] Determine a training set based on the historical image set; and train the untrained YOLOv8 defect detection model based on the training set to obtain a trained YOLOv8 defect detection model;

[0008] Obtain real-time image data of the UAV during inspection of the target power area, and determine defect information about the target power area based on the trained YOLOv8 defect detection model and the real-time image data.

[0009] In some possible embodiments, the obtaining the historical image set of the UAV during inspection of the target power area includes:

[0010] Obtain an initial historical image set of the UAV during inspection of the target power area; wherein, the initial historical image set includes multiple initial historical images about the target power area;

[0011] Add defect labels to each initial historical image in the initial historical image set based on a preset defect standard to obtain an intermediate historical image set; wherein, the intermediate historical image set includes multiple labeled intermediate historical images, and the defect labels include defect positions, defect types, and defect grades;

[0012] Perform linear interpolation on the multiple labeled intermediate historical images in the intermediate historical image set based on data augmentation techniques, and determine the historical image set based on the linear interpolation results and the intermediate historical image set.

[0013] In some possible embodiments, the training set includes a training data set and a training label set; the training the YOLOv8 defect detection model to be trained based on the training set includes:

[0014] Obtain the YOLOv8 defect detection model to be trained; wherein, the YOLOv8 defect detection model to be trained includes a feature extraction module, a feature fusion module, and a defect prediction module;

[0015] For each training data in the training data set, perform feature extraction on the training data based on the feature extraction module to obtain feature information of different scales corresponding to the training data; and, perform feature fusion processing on the feature information of different scales based on the feature fusion module to obtain feature fusion information; and, perform prediction on the feature fusion information based on the defect prediction module to obtain prediction defect information corresponding to the training data;

[0016] Calculate a loss value between the prediction defect information corresponding to the training data and the defect label of the training sample data corresponding to the training data in the training label set based on a distance loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value;

[0017] Repeat the above steps until the training result meets the preset requirements to obtain the trained YOLOv8 segmentation model.

[0018] In some possible embodiments, the feature extraction module includes a global feature extraction module, a RepResNet feature extraction module, and a feature initial fusion module. The global feature extraction module includes a dilated convolution module and a common convolution module; performing feature extraction on the training data based on the feature extraction module includes:

[0019] Performing feature extraction processing on the training data based on the dilated convolution module to obtain first feature information; and performing feature extraction processing on the training data based on the common convolution module to obtain second feature information; and determining a first feature map based on the first feature information and the second feature information; and performing multi-scale feature enhancement processing on the first feature map based on the RepResNet feature extraction module to determine a first feature representation; and determining first-scale feature information based on the feature initial fusion module, the first feature map, and the first feature representation;

[0020] Performing feature extraction processing on the first-scale feature information based on the dilated convolution module to obtain third feature information; and performing feature extraction processing on the first-scale feature information based on the common convolution module to obtain fourth feature information; and determining a second feature map based on the third feature information and the fourth feature information; and performing multi-scale feature enhancement processing on the second feature map based on the RepResNet feature extraction module to determine a second feature representation; and determining second-scale feature information based on the feature initial fusion module, the second feature map, and the second feature representation;

[0021] Performing feature extraction processing on the second-scale feature information based on the dilated convolution module to obtain fifth feature information; and performing feature extraction processing on the second-scale feature information based on the common convolution module to obtain sixth feature information; and determining a third feature map based on the fifth feature information and the sixth feature information; and performing multi-scale feature enhancement processing on the third feature map based on the RepResNet feature extraction module to determine a third feature representation; and determining third-scale feature information based on the feature initial fusion module, the third feature map, and the third feature representation;

[0022] Performing feature extraction processing on the third-scale feature information based on the dilated convolution module to obtain seventh feature information; performing feature extraction processing on the third-scale feature information based on the ordinary convolution module to obtain eighth feature information; and determining a fourth feature map based on the seventh feature information and the eighth feature information; and performing multi-scale feature enhancement processing on the fourth feature map based on the RepResNet feature extraction module to determine a fourth feature representation; and determining fourth-scale feature information based on the feature initial fusion module, the fourth feature map, and the fourth feature representation.

[0023] In some possible embodiments, the performing feature fusion processing on the feature information of different scales based on the feature fusion module includes:

[0024] Performing feature fusion processing on the second-scale feature information, the third-scale feature information, and the fourth-scale feature information based on the feature fusion module to obtain the feature fusion information.

[0025] In some possible embodiments, the distance loss function includes:

[0026]

[0027] Where, L represents the loss value; IoU represents the degree of overlap between the position box of the predicted defect and the position box of the defect label; Distance represents the Euclidean distance metric between the center point of the position box of the predicted defect and the center point of the position box of the defect label.

[0028] In some possible embodiments, after determining the defect information about the target power region based on the trained YOLOv8 defect detection model and the real-time image data, it further includes:

[0029] Determining the defect cause and repair suggestion according to the preset defect knowledge base, the defect information, and the real-time image data.

[0030] The embodiments of the present disclosure provide an unmanned aerial vehicle inspection device for online detection of a power system, including:

[0031] An image acquisition module, configured to acquire a historical image set of the unmanned aerial vehicle in the inspected target power region; wherein, the historical image set includes multiple labeled historical images about the target power region;

[0032] A model training module, configured to determine a training set based on the historical image set; and train a YOLOv8 defect detection model to be trained based on the training set to obtain a trained YOLOv8 defect detection model;

[0033] A defect determination module, configured to obtain real-time image data of the UAV during inspection of the target power area, and determine defect information about the target power area based on the trained YOLOv8 defect detection model and the real-time image data.

[0034] In some possible embodiments, the image acquisition module is specifically configured to:

[0035] Obtain an initial historical image set of the UAV during inspection of the target power area; wherein, the initial historical image set includes multiple initial historical images about the target power area;

[0036] Add defect labels to each initial historical image in the initial historical image set based on a preset defect standard, to obtain an intermediate historical image set; wherein, the intermediate historical image set includes multiple labeled intermediate historical images, and the defect labels include defect positions, defect types, and defect levels;

[0037] Perform linear interpolation on the multiple labeled intermediate historical images in the intermediate historical image set based on data augmentation techniques, and determine the historical image set based on the linear interpolation result and the intermediate historical image set.

[0038] In some possible embodiments, the model training module is specifically configured to:

[0039] Obtain the YOLOv8 defect detection model to be trained; wherein, the YOLOv8 defect detection model to be trained includes a feature extraction module, a feature fusion module, and a defect prediction module;

[0040] For each training data in the training dataset, perform feature extraction on the training data based on the feature extraction module to obtain feature information of different scales corresponding to the training data; and, perform feature fusion processing on the feature information of different scales based on the feature fusion module to obtain feature fusion information; and, perform prediction on the feature fusion information based on the defect prediction module to obtain prediction defect information corresponding to the training data;

[0041] Calculate a loss value between the prediction defect information corresponding to the training data and the defect labels of the training sample data corresponding to the training data in the training label set based on a distance loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value;

[0042] Repeat the above steps until the training result meets the preset requirements, to obtain the trained YOLOv8 segmentation model.

[0043] In some possible embodiments, the feature extraction module includes a global feature extraction module, a RepResNet feature extraction module, and a feature preliminary fusion module. The global feature extraction module includes a dilated convolution module and a common convolution module. The model training module is further configured to:

[0044] perform feature extraction processing on the training data based on the dilated convolution module to obtain first feature information; perform feature extraction processing on the training data based on the common convolution module to obtain second feature information; and determine a first feature map based on the first feature information and the second feature information; and perform multi-scale feature enhancement processing on the first feature map based on the RepResNet feature extraction module to determine a first feature representation; and determine first-scale feature information based on the feature preliminary fusion module, the first feature map, and the first feature representation;

[0045] perform feature extraction processing on the first-scale feature information based on the dilated convolution module to obtain third feature information; perform feature extraction processing on the first-scale feature information based on the common convolution module to obtain fourth feature information; and determine a second feature map based on the third feature information and the fourth feature information; and perform multi-scale feature enhancement processing on the second feature map based on the RepResNet feature extraction module to determine a second feature representation; and determine second-scale feature information based on the feature preliminary fusion module, the second feature map, and the second feature representation;

[0046] perform feature extraction processing on the second-scale feature information based on the dilated convolution module to obtain fifth feature information; perform feature extraction processing on the second-scale feature information based on the common convolution module to obtain sixth feature information; and determine a third feature map based on the fifth feature information and the sixth feature information; and perform multi-scale feature enhancement processing on the third feature map based on the RepResNet feature extraction module to determine a third feature representation; and determine third-scale feature information based on the feature preliminary fusion module, the third feature map, and the third feature representation;

[0047] perform feature extraction processing on the third-scale feature information based on the dilated convolution module to obtain seventh feature information; perform feature extraction processing on the third-scale feature information based on the common convolution module to obtain eighth feature information; and determine a fourth feature map based on the seventh feature information and the eighth feature information; and perform multi-scale feature enhancement processing on the fourth feature map based on the RepResNet feature extraction module to determine a fourth feature representation; and determine fourth-scale feature information based on the feature preliminary fusion module, the fourth feature map, and the fourth feature representation.

[0048] In some possible embodiments, the model training module is further configured to:

[0049] Based on the feature fusion module, perform feature fusion processing on the second-scale feature information, the third-scale feature information, and the fourth-scale feature information to obtain the feature fusion information.

[0050] In some possible embodiments, the distance loss function includes:

[0051]

[0052] where L represents the loss value; IoU represents the degree of overlap between the position box of the predicted defect and the position box of the defect label; Distance represents the Euclidean distance metric between the center point of the position box of the predicted defect and the center point of the position box of the defect label.

[0053] In some possible embodiments, the defect determination module is further configured to:

[0054] Determine the defect cause and repair suggestions according to the preset defect knowledge base, the defect information, and the real-time image data.

[0055] An embodiment of the present disclosure provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the unmanned aerial vehicle inspection method for online detection of power systems described in any of the above possible implementation manners is executed.

[0056] An embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the unmanned aerial vehicle inspection method for online detection of power systems described in any of the above possible implementation manners is implemented.

[0057] The UAV inspection method, device, and medium for online detection of power systems provided in the embodiments of the present disclosure specifically first obtain a historical image set of the UAV in the target power area to be inspected; among them, the historical image set includes multiple labeled historical images with defect annotations of the target power area. Based on these labeled historical images, a training set for model training can be further determined; subsequently, the YOLOv8 defect detection model to be trained is trained based on the training set, and through multiple iterations and optimizations, a trained YOLOv8 defect detection model is obtained. Next, real-time image data of the UAV in the target power area to be inspected is obtained and input into the trained YOLOv8 defect detection model, and defect information about the target power area, such as defect type, location, and severity, is determined based on the output of the model.

[0058] In this way, the present disclosure trains a high-precision YOLOv8 defect detection model by using a historical image set with defect annotations, and applies this model for defect detection in real-time inspections, achieving fast response and accurate identification of online detection of power systems, effectively improving the efficiency and accuracy of inspections, and reducing the costs and risks of manual inspections.

[0059] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be cited in the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0061] Figure 1 Shows a flowchart of a UAV inspection method for online detection of power systems provided in an embodiment of the present disclosure;

[0062] Figure 2 Shows a flowchart of a model training method provided in an embodiment of the present disclosure;

[0063] Figure 3 Shows a flowchart of a feature extraction method provided in an embodiment of the present disclosure;

[0064] Figure 4 Shows a schematic diagram of a module for global feature extraction provided in an embodiment of the present disclosure;

[0065] Figure 5 shows a schematic structural diagram of an unmanned aerial vehicle inspection device for online detection of a power system provided by an embodiment of the present disclosure;

[0066] Figure 6 shows a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Components of the embodiments of the present disclosure usually described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0068] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0069] The term "and / or" in this article merely describes an associative relationship and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, both A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this article means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.

[0070] The power system is an indispensable infrastructure in modern society and has important practical significance for ensuring power supply safety and stability. However, there are various potential unknown faults and operation defects in the power system, such as insulator breakage, transmission line looseness, monitoring equipment aging, etc. If these problems cannot be discovered and handled in time, power accidents such as power outages may occur, resulting in serious consequences. Traditional power system inspection methods rely on manual inspection and regular maintenance. Manual inspection requires a large amount of human and time costs, has low efficiency and subjective judgment biases; at the same time, the areas where power infrastructure is located are vast and complex, and there are often high, inaccessible or dangerous areas, which pose safety hazards to manual inspection.

[0071] It has been found through research that with the rapid development of drone technology and edge computing, online inspection can be carried out by taking advantage of the characteristics of drones flying flexibly over various areas of the power system, without the need for personnel to directly enter dangerous areas, overcoming various limitations of traditional inspection methods. However, the defect detection methods in related technologies mostly rely on uploading image data to the cloud for analysis and processing. This method has problems such as high data transmission latency and large network bandwidth requirements, and cannot respond to fault warnings in real time; moreover, there may be privacy leakage and security issues during the data transmission process.

[0072] Based on the above research, in the embodiments of the present disclosure, a drone inspection method, device, and medium for online detection of power systems are provided. Specifically, first, a historical image set of the drone in the target power area to be inspected is obtained; among them, the historical image set includes multiple labeled historical images with defect annotations regarding the target power area. Based on these labeled historical images, a training set for model training can be further determined; subsequently, the YOLOv8 defect detection model to be trained is trained based on the training set, and through multiple iterations and optimizations, a trained YOLOv8 defect detection model is obtained. Next, real-time image data of the drone in the target power area to be inspected is obtained and input into the trained YOLOv8 defect detection model, and defect information regarding the target power area, such as defect type, location, and severity, is determined based on the output of the model.

[0073] In the embodiments of the present disclosure, a high-precision YOLOv8 defect detection model is trained by using a historical image set with defect annotations, and this model is applied for defect detection during real-time inspection, achieving rapid response and accurate identification for online detection of power systems, effectively improving the efficiency and accuracy of inspection, and reducing the cost and risk of manual inspection.

[0074] To facilitate the understanding of this embodiment, first, the execution subject of the drone inspection method for online detection of power systems provided by the embodiments of the present disclosure is introduced in detail. The execution subject of the drone inspection method for online detection of power systems provided by the embodiments of the present disclosure is a computer device. This computer device can be a server. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms.

[0075] The following will detail the drone inspection method for online detection of power systems provided by the embodiments of the present application with reference to the accompanying drawings. See Figure 1As shown in the figure, it is a flowchart of an unmanned aerial vehicle (UAV) inspection method for online detection of a power system provided by an embodiment of the present disclosure. The UAV inspection method for online detection of a power system includes the following S101 to S103:

[0076] S101, obtain a historical image set of the UAV in the target power area to be inspected.

[0077] It can be understood that the historical image set refers to a set of images captured and stored by the UAV when performing inspection tasks on the target power area in the past, and may include multiple labeled historical images of the target power area.

[0078] Specifically, in the process of obtaining the historical image set of the target power area, the following steps (1) to (3) may be included:

[0079] (1) Obtain an initial historical image set of the UAV in the target power area to be inspected; wherein, the initial historical image set includes multiple initial historical images of the target power area;

[0080] (2) Add defect labels to each initial historical image in the initial historical image set based on a preset defect standard to obtain an intermediate historical image set; wherein, the intermediate historical image set includes multiple labeled intermediate historical images, and the defect labels include defect positions, defect types, and defect levels;

[0081] (3) Perform linear interpolation on the multiple labeled intermediate historical images in the intermediate historical image set based on data augmentation technology, and determine the historical image set based on the linear interpolation result and the intermediate historical image set.

[0082] It can be understood that when the UAV performs inspection tasks, it will take all-round pictures of the target power area according to a preset flight path and shooting parameters. These captured images, after being screened and sorted, constitute the initial historical image set. These images may contain views of various power facilities (such as power towers, wires, substations, etc.) and their states at different time periods and different weather conditions. Further, after obtaining the initial historical image set, these images need to be labeled, mainly according to the standard defects of power equipment, that is, defect labels. The purpose of labeling is to identify and mark the defect information in the images according to the preset defect standard. This includes the position of the defect (such as the break point of the wire, the corroded part of the power tower, etc.), the type of the defect (such as breakage, corrosion, missing, etc.), and the level of the defect (such as minor, medium, severe, etc.). The labeling process can be completed with the help of professional labeling tools (such as labelme) to ensure the accuracy and consistency of labeling. The labeled images form an intermediate historical image set, which provides labeled training samples for subsequent model training.

[0083] In some possible embodiments, before annotating the defect information of an image, a specification library can also be established to standardize the format requirements and defect types of pre-annotated data, and regularly maintain the annotation standards. The annotation library is extended according to new defect types and annotation requirements to better evaluate the defect status. The format requirements mainly include the accurate annotation of the positions of defect targets in the image and the drawing criteria for defect bounding boxes; the defect types mainly determine the types of target defects to be detected and annotated, such as cracks, corrosion, deformation, etc. At the same time, the severity levels of different types of defects are defined, such as mild, medium, severe levels. The defect information of each image is annotated through the specification library.

[0084] In some other embodiments, after obtaining the initial historical image set, image preprocessing operations such as denoising can also be performed on the images collected therein, which are not specifically limited herein.

[0085] Here, in order to enhance the generalization ability of the model, the present disclosure uses data augmentation techniques. Data augmentation refers to generating more image samples by operating on existing images (such as rotation, cropping, scaling, etc.). Among them, linear interpolation is a common image enhancement method, and new image data can be generated through existing images. After enhancing the annotated intermediate historical image set through linear interpolation technology, a final historical image set is formed. The images processed by data augmentation not only retain the key information in the original images, but also increase the diversity and complexity of the images. These enhanced images, together with the original annotated images, constitute the final historical image set. This set provides richer and more diverse training samples for subsequent model training and verification.

[0086] Specifically, linear interpolation can be expressed as:

[0087]

[0088] Among them, represents the new sample data pair obtained by interpolation; (A i , B i ) represents the i-th sample data pair of the intermediate historical image set; (A j , B j ) represents the j-th sample data pair of the intermediate historical image set; A represents the intermediate historical image; B represents the label corresponding to the intermediate historical image; λ represents a random scaling factor used to control the interpolation weight between sample A and sample B.

[0089] S102. Determine a training set based on the historical image set; and train the YOLOv8 defect detection model to be trained based on the training set to obtain a trained YOLOv8 defect detection model.

[0090] It can be understood that the YOLOv8 defect detection model is an object detection algorithm widely used in object recognition and localization in images. In the embodiments of the present disclosure, YOLOv8 is used to detect various defects in power facilities. Through the training of the historical image set, the YOLOv8 model can learn how to identify the defects of power facilities from images.

[0091] Specifically, the training set includes a training data set and a training label set. Referring to Figure 2 shown, a model training method proposed by the present disclosure may include the following S201 to S204:

[0092] S201, obtain the YOLOv8 defect detection model to be trained.

[0093] Among them, the YOLOv8 defect detection model to be trained includes a feature extraction module, a feature fusion module, and a defect prediction module; the feature extraction module is responsible for extracting feature information of different scales from the input image. These feature information contain key information such as the shape and texture of the objects in the image. The feature fusion module then fuses these feature information of different scales to generate a more robust and comprehensive feature representation. Here, the feature extraction module includes a global feature extraction module, a RepResNet feature extraction module, and a feature initial fusion module. The global feature extraction module includes a dilated convolution module and a common convolution module. Finally, the defect prediction module predicts the defects in the image based on the fused feature information and outputs the predicted defect information.

[0094] S202, for each training data in the training data set, based on the feature extraction module, extract the feature information of different scales corresponding to the training data; and, based on the feature fusion module, perform feature fusion processing on the feature information of different scales to obtain feature fusion information; and, based on the defect prediction module, predict the feature fusion information to obtain the predicted defect information corresponding to the training data.

[0095] Specifically, for the data in each training data set, first, the image is processed by the feature extraction module to extract the feature information of different scales corresponding to the training data. These feature information provide rich descriptions in terms of the position, shape, texture, etc. of the image in the image, helping the model understand the image content. Then, using the feature fusion module, the feature information of different scales is fused to form a feature representation with more global information. Through this multi-scale feature fusion, the model can better capture the details of the defects in the power facilities and enhance the model's detection ability for defects of different sizes. Finally, based on the defect prediction module, using the fused feature information, the defects are predicted, thereby obtaining the predicted defect information corresponding to each training data.

[0096] Exemplarily, in order to achieve efficient feature extraction and enhancement, when processing the training data based on the feature extraction module, the understanding and representation ability of the model for the data can be gradually improved through iterative calculations in multiple stages. Refer to Figure 3 As shown, when performing feature extraction on the training data based on the feature extraction module, the following S2021 to S2024 may be included:

[0097] S2021, perform feature extraction processing on the training data based on the dilated convolution module to obtain first feature information; and perform feature extraction processing on the training data based on the ordinary convolution module to obtain second feature information; and, determine a first feature map based on the first feature information and the second feature information; and, perform multi-scale feature enhancement processing on the first feature map based on the RepResNet feature extraction module to determine a first feature representation; and determine first-scale feature information based on the feature initial fusion module, the first feature map, and the first feature representation.

[0098] Specifically, first, the dilated convolution module is used to process the training data. The dilated convolution can increase the receptive field of the convolution kernel without increasing the computational complexity, thereby capturing more extensive context information to obtain the first feature information. At the same time, the ordinary convolution module is also used to extract more local and finer features to generate the second feature information. These two feature information are then integrated to form a first feature map (i.e., the output shown in the figure), which contains both global information and local details. Here, refer to Figure 4 As shown, this is a schematic diagram of a global feature extraction module proposed by the present disclosure. The original input (i.e., the training data) can be divided into two parts. One part is input into a 3×3 dilated convolution for context feature extraction, and the other part is input into a 3×3 ordinary convolution for conventional extraction. Feature concatenation is performed along the channel dimension, and after normalization and activation function processing, a first feature map (i.e., the output part) is obtained for input to the next module.

[0099] Furthermore, due to the increase in network depth, ordinary convolutional networks are prone to problems such as vanishing gradients or exploding gradients, making it difficult to train the network. By introducing the RepResNet structure with residual connections and dense connections, which includes three convolutional layers (Branch 1: 3×3 convolution, Branch 2: 1×1 convolution, Branch 3: direct connection) and four cascaded and stacked modules, it can enhance the ability of feature reuse and information transmission. Through the RepResNet feature extraction module (an improved deep residual network), multi-scale feature enhancement processing is performed on the first feature map. Different-scale convolutional kernels and residual connections are used to enhance the expressive power and robustness of the feature map, obtaining the first feature representation. Finally, the feature initial fusion module fuses the first feature map and the first feature representation to determine the first-scale feature information. Among them, the feature initial fusion module can be implemented using the C2F structure (i.e., Cross-scale Feature Fusion). The C2F structure can effectively fuse feature information from different scales and levels, ensuring that the network can capture multi-scale and multi-level feature representations.

[0100] In the embodiments of the present disclosure, by introducing the RepResNet structure and the C2F feature fusion module, the entire network can effectively solve the common problems of vanishing gradients and exploding gradients in deep networks and enhance the expressive power and robustness of the feature map. The combination of residual connections and dense connections makes the network perform more efficiently in information transmission and feature reuse, ultimately improving the performance of the network in processing multi-scale and complex visual tasks. This improved deep residual network demonstrates strong capabilities in feature extraction and multi-scale feature enhancement, and can help the network achieve better results in various image recognition tasks.

[0101] In some possible embodiments, the ordinary convolutional layers in the backbone network can be replaced by using CGConv, aiming to enhance the receptive field of the initial features and the context modeling ability through learnable complementary feature patterns, which will not be specifically limited here.

[0102] S2022, perform feature extraction processing on the first-scale feature information based on the dilated convolution module to obtain the third feature information; and perform feature extraction processing on the first-scale feature information based on the ordinary convolution module to obtain the fourth feature information; and determine the second feature map based on the third feature information and the fourth feature information; and perform multi-scale feature enhancement processing on the second feature map based on the RepResNet feature extraction module to determine the second feature representation; and determine the second-scale feature information based on the feature initial fusion module, the second feature map, and the second feature representation.

[0103] Here, the process of the above step S2022 is repeatedly applied to the first-scale feature information. The dilated convolution module and the ordinary convolution module respectively extract the third feature information and the fourth feature information again, and integrate them into the second feature map. Subsequently, the RepResNet feature extraction module performs multi-scale feature enhancement on the second feature map to generate the second feature representation. The feature initial fusion module plays a role again, fusing the second feature map with the second feature representation to obtain the second-scale feature information. In this way, the model can gradually and deeply mine the multi-level features in the data.

[0104] S2023, perform feature extraction processing on the second-scale feature information based on the dilated convolution module to obtain the fifth feature information; and perform feature extraction processing on the second-scale feature information based on the ordinary convolution module to obtain the sixth feature information; and, determine the third feature map based on the fifth feature information and the sixth feature information; and, perform multi-scale feature enhancement processing on the third feature map based on the RepResNet feature extraction module to determine the third feature representation; and determine the third-scale feature information based on the feature initial fusion module, the third feature map, and the third feature representation.

[0105] Here, the above step S2022 is further processed on the second-scale feature information as well. The dilated convolution module and the ordinary convolution module respectively extract the fifth feature information and the sixth feature information to form the third feature map. Through the multi-scale feature enhancement of the RepResNet feature extraction module, the third feature representation is obtained. The feature initial fusion module fuses the third feature map with the third feature representation to generate the third-scale feature information, enabling the model to gradually capture deeper-level features in the data.

[0106] S2024, perform feature extraction processing on the third-scale feature information based on the dilated convolution module to obtain the seventh feature information; and perform feature extraction processing on the third-scale feature information based on the ordinary convolution module to obtain the eighth feature information; and, determine the fourth feature map based on the seventh feature information and the eighth feature information; and, perform multi-scale feature enhancement processing on the fourth feature map based on the RepResNet feature extraction module to determine the fourth feature representation; and determine the fourth-scale feature information based on the feature initial fusion module, the fourth feature map, and the fourth feature representation.

[0107] Specifically, the dilated convolution module and the ordinary convolution module respectively extract the seventh feature information and the eighth feature information to form the fourth feature map. The RepResNet feature extraction module performs multi-scale feature enhancement on the fourth feature map to generate the fourth feature representation. Finally, the feature initial fusion module fuses the fourth feature map and the fourth feature representation to obtain the fourth-scale feature information. At this point, the model has gone through four iterative processes and gradually and deeply extracted feature representations containing different scales and different levels of information from the training data.

[0108] In the embodiments of the present disclosure, based on the trade-off between model performance and computational complexity, the present disclosure adopts a method of performing four iterations on the training data to achieve the extraction of feature information of different scales. Here, the number of iterations is not fixed. In some other embodiments, if the dimension of the training data is relatively high or the complexity is relatively large, in order to more fully capture the multi-level features in the data, five, six or more iterative processes can be used. By increasing the number of iterations, the model can learn richer and more detailed feature representations, thereby potentially improving the performance on certain specific tasks. At the same time, it is also necessary to balance the computational overhead and time cost brought by the increase in the number of iterations. In practical applications, the appropriate number of iterations should be selected according to specific requirements and resource limitations, and no specific limitation is made here.

[0109] It can be understood that, in order to effectively utilize the complementarity between different-scale feature information and reduce redundant information, so as to achieve higher accuracy and robustness in the object detection task, after obtaining the first-scale feature information, the second-scale feature information, the third-scale feature information, and the fourth-scale feature information, the present disclosure does not directly fuse all-scale feature information. Instead, when using the feature fusion module to perform feature fusion processing on the different-scale feature information, the feature fusion module mainly performs feature fusion processing on the second-scale feature information, the third-scale feature information, and the fourth-scale feature information. Generally, deeper-level features (such as the third-scale and fourth-scale feature information) contain higher-level semantic information and broader context awareness, which are crucial for category recognition and precise location in object detection. Although the shallower-level features (such as the first-scale feature information) contain rich detailed information, after multiple iterations and feature enhancements, some of their information may have been covered or surpassed by deeper-level features.

[0110] Therefore, in the feature fusion stage, by ignoring or reducing the dependence on the first-scale feature information and focusing more on the fusion of the second-scale, third-scale, and fourth-scale feature information, the present disclosure can more effectively integrate the feature advantages of different scales while reducing unnecessary computational overhead.

[0111] S203. Calculate the loss value between the predicted defect information corresponding to the training data and the defect label of the training sample data corresponding to the training data in the training label set based on the distance loss function, and adjust the to-be-trained YOLOv8 segmentation model based on the loss value.

[0112] It can be understood that the distance loss function is used to calculate the difference between the predicted defect information and the true defect label in the training label set. The loss function quantifies the error in the prediction by measuring the gap between the predicted value and the true label. By calculating the loss value, the model can evaluate its prediction performance and thus be optimized. This loss value reflects the deviation between the prediction and the true defect label, and the model will adjust its internal parameters (such as weights and biases) according to this loss value to better match the defect annotations in the training data. This process can continuously adjust the model parameters through optimization algorithms such as backpropagation and gradient descent until the model can achieve good prediction accuracy on the training data.

[0113] Here, in order to further enhance the performance of the improved YOLOv8 object detection model, the present disclosure introduces the WIoUv1 distance loss function to replace the regression loss CIOU in the original model. The CIOU loss optimizes the positioning accuracy of the bounding box, but still has deficiencies such as high sensitivity, inapplicability to incomplete bounding boxes, insensitivity to aspect ratios, and limitation to the regression range. The WIoUv1 introduces a distance attention mechanism, takes the positioning accuracy into weighted consideration, and adaptively weights according to the distance of the target, thus more flexibly adjusting the weight of the positioning accuracy and more comprehensively optimizing the positioning of the bounding box. Its loss function can be expressed as:

[0114]

[0115] Among them, L represents the loss value; IoU represents the overlapping degree between the position box of the predicted defect and the position box of the defect label; Distance represents the Euclidean distance metric between the center points of the position boxes of the predicted defect and the defect label.

[0116] S204: Repeat the above steps until the training result meets the preset requirements, and obtain the trained YOLOv8 segmentation model.

[0117] Specifically, the above steps S201 to S204 will be repeated multiple times until the training result meets the preset accuracy requirements. Each iteration will be optimized and adjusted based on the loss value of the previous step, gradually improving the prediction accuracy of the model. After multiple iterations of training, the model will gradually learn how to more accurately identify defects in power facilities and finally generate a trained YOLOv8 defect detection model.

[0118] S103. Obtain the real-time image data of the UAV during the inspection of the target power area, and determine the defect information about the target power area based on the trained YOLOv8 defect detection model and the real-time image data.

[0119] Specifically, deploy the fully trained and verified YOLOv8 defect detection model to the inspection system of the UAV. The UAV uses its equipped high-definition camera to capture the image data of the target power area in real time. These image data are then input into the trained YOLOv8 defect detection model, which can accurately detect and locate potential defects in the images. Furthermore, it outputs the category of the defect (such as line breakage, equipment damage, etc.), the precise position of the defect in the image, and the level of the defect, thus achieving real-time target defect detection.

[0120] In some possible embodiments, after the model identifies the defect target, the defect cause and repair suggestions can be determined based on the preset defect knowledge base, defect information, and real-time image data; a detailed defect result report can also be further generated, including information such as the accuracy evaluation of the defect, specific category, severity, etc. These information are immediately fed back to the manual operators at the client side so that they can take corresponding handling measures according to the actual situation of the defect. The handling methods may include immediately notifying relevant personnel to go to the site to repair the defect, marking the defect location in the system for subsequent tracking and handling, or formulating a long-term maintenance plan according to the severity of the defect.

[0121] In some possible embodiments, in order to continuously improve the performance and accuracy of the detection model, the data and results of defect handling can also be collected and recorded. These data cover key information such as the category of the defect, location marking, severity, and maintenance records. By deeply analyzing these data, the detection model can be further optimized, which may include adding new training data to cover more types of defects, improving the quality of data annotation to ensure the accuracy of the model, adjusting the hyperparameters of the model to optimize its performance, or introducing new technologies or architectures to enhance the detection ability of the model, so as to continuously improve the accuracy and stability of defect detection and provide more reliable technical support for power inspection.

[0122] The UAV inspection method, device, and medium provided in the embodiments of the present disclosure for online detection of power systems train a high-precision YOLOv8 defect detection model by using a historical image set with defect annotations, and apply this model for defect detection in real-time inspections, achieving fast response and accurate identification for online detection of power systems, effectively improving the efficiency and accuracy of inspections, and reducing the costs and risks of manual inspections.

[0123] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0124] Based on the same inventive concept, an unmanned aerial vehicle (UAV) inspection device for power system online detection corresponding to the UAV inspection method for power system online detection is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above UAV inspection method for power system online detection in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0125] Refer to Figure 5 As shown, it is a schematic diagram of an unmanned aerial vehicle (UAV) inspection device 500 for power system online detection provided by an embodiment of the present disclosure. The device includes:

[0126] An image acquisition module 501, configured to acquire a historical image set of the UAV in the target power area to be inspected; wherein, the historical image set includes multiple labeled historical images of the target power area;

[0127] A model training module 502, configured to determine a training set based on the historical image set; and train a YOLOv8 defect detection model to be trained based on the training set to obtain a trained YOLOv8 defect detection model;

[0128] A defect determination module 503, configured to acquire real-time image data of the UAV in the target power area to be inspected, and determine defect information about the target power area based on the trained YOLOv8 defect detection model and the real-time image data.

[0129] In some possible embodiments, the image acquisition module 501 is specifically configured to:

[0130] Acquire an initial historical image set of the UAV in the target power area to be inspected; wherein, the initial historical image set includes multiple initial historical images of the target power area;

[0131] Add defect labels to each initial historical image in the initial historical image set based on a preset defect standard to obtain an intermediate historical image set; wherein, the intermediate historical image set includes multiple labeled intermediate historical images, and the defect labels include defect positions, defect types, and defect levels;

[0132] Perform linear interpolation on the multiple labeled intermediate historical images in the intermediate historical image set based on data augmentation technology, and determine the historical image set based on the linear interpolation result and the intermediate historical image set.

[0133] In some possible embodiments, the model training module 502 is specifically configured to:

[0134] Obtain the YOLOv8 defect detection model to be trained; wherein, the YOLOv8 defect detection model to be trained includes a feature extraction module, a feature fusion module, and a defect prediction module;

[0135] For each piece of training data in the training dataset, perform feature extraction on the training data based on the feature extraction module to obtain feature information of different scales corresponding to the training data; and, perform feature fusion processing on the feature information of different scales based on the feature fusion module to obtain feature fusion information; and, perform prediction on the feature fusion information based on the defect prediction module to obtain prediction defect information corresponding to the training data;

[0136] Calculate the loss value between the prediction defect information corresponding to the training data and the defect label of the training sample data corresponding to the training data in the training label set based on the distance loss function, and adjust the YOLOv8 segmentation model to be trained based on the loss value;

[0137] Repeat the above steps until the training result meets the preset requirements to obtain the trained YOLOv8 segmentation model.

[0138] In some possible embodiments, the feature extraction module includes a global feature extraction module, a RepResNet feature extraction module, and a feature initial fusion module, and the global feature extraction module includes a dilated convolution module and a common convolution module; the model training module 502 is further configured to:

[0139] Perform feature extraction processing on the training data based on the dilated convolution module to obtain first feature information; and perform feature extraction processing on the training data based on the common convolution module to obtain second feature information; and, determine a first feature map based on the first feature information and the second feature information; and, perform multi-scale feature enhancement processing on the first feature map based on the RepResNet feature extraction module to determine a first feature representation; and determine first-scale feature information based on the feature initial fusion module, the first feature map, and the first feature representation;

[0140] Performing feature extraction processing on the first-scale feature information based on the dilated convolution module to obtain third feature information; performing feature extraction processing on the first-scale feature information based on the ordinary convolution module to obtain fourth feature information; and determining a second feature map based on the third feature information and the fourth feature information; and performing multi-scale feature enhancement processing on the second feature map based on the RepResNet feature extraction module to determine a second feature representation; and determining second-scale feature information based on the feature initial fusion module, the second feature map, and the second feature representation;

[0141] Performing feature extraction processing on the second-scale feature information based on the dilated convolution module to obtain fifth feature information; performing feature extraction processing on the second-scale feature information based on the ordinary convolution module to obtain sixth feature information; and determining a third feature map based on the fifth feature information and the sixth feature information; and performing multi-scale feature enhancement processing on the third feature map based on the RepResNet feature extraction module to determine a third feature representation; and determining third-scale feature information based on the feature initial fusion module, the third feature map, and the third feature representation;

[0142] Performing feature extraction processing on the third-scale feature information based on the dilated convolution module to obtain seventh feature information; performing feature extraction processing on the third-scale feature information based on the ordinary convolution module to obtain eighth feature information; and determining a fourth feature map based on the seventh feature information and the eighth feature information; and performing multi-scale feature enhancement processing on the fourth feature map based on the RepResNet feature extraction module to determine a fourth feature representation; and determining fourth-scale feature information based on the feature initial fusion module, the fourth feature map, and the fourth feature representation.

[0143] In some possible embodiments, the model training module 502 is further configured to:

[0144] Performing feature fusion processing on the second-scale feature information, the third-scale feature information, and the fourth-scale feature information based on the feature fusion module to obtain the feature fusion information.

[0145] In some possible embodiments, the distance loss function includes:

[0146]

[0147] where L represents the loss value; IoU represents the degree of overlap between the position box of the predicted defect and the position box of the defect label; Distance represents the Euclidean distance metric between the center points of the position boxes of the predicted defect and the position box of the defect label.

[0148] In some possible embodiments, the defect determination module 503 is further configured to:

[0149] Determine the defect cause and repair suggestions according to a preset defect knowledge base, the defect information, and the real-time image data.

[0150] Based on the same inventive concept, an embodiment of the present disclosure further provides a computer device. Referring to Figure 6 As shown, it is a schematic structural diagram of a computer device 600 provided by an embodiment of the present disclosure, including a processor 601, a memory 602, and a bus 603. Among them, the memory 602 is used to store execution instructions, including an internal memory 6021 and an external memory 6022; the internal memory 6021 here is also called the main memory, which is used to temporarily store the operation data in the processor 601 and the data exchanged with the external memory 6022 such as a hard disk. The processor 601 exchanges data with the external memory 6022 through the internal memory 6021. <??

[0151] In an embodiment of the present application, the memory 602 is specifically configured to store the application program code for implementing the solution of the present application, and is controlled by the processor 601 to execute. That is, when the computer device 600 runs, the processor??601 communicates with the memory 602 through the bus 603, so that the processor 601 executes the application program code stored in the memory 602, and further executes the method described in any one of the foregoing embodiments.

[0152] Among them, the memory 602 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0153] It should be noted that there seems to be an error in line 12 where "processor??601" is an incorrect expression. It should probably be "processor 601".The processor 601 may be an integrated circuit chip with the ability to process signals. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0154] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the computer device 600. In other embodiments of the present application, the computer device 600 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0155] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the drone inspection method for online detection of power systems described in the above method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0156] The embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the drone inspection method for online detection of power systems described in the above method embodiments. For details, refer to the above method embodiments and will not be elaborated here.

[0157] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0158] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0159] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0160] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0161] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0162] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A drone inspection method for online detection of power systems, characterized in that: include: Acquire a historical image set of the drone during inspection of a target power area; wherein the historical image set includes a plurality of labeled historical images of the target power area; Determining a training set based on the historical image set; and training a YOLOv8 defect detection model to be trained based on the training set to obtain a trained YOLOv8 defect detection model; Real-time image data of the target power area being inspected by the drone is obtained, and defect information about the target power area is determined based on the trained YOLOv8 defect detection model and the real-time image data.

2. The method according to claim 1, characterized in that The acquisition of a historical image set of the drone inspecting the target power area includes: Acquire an initial historical image set of the drone during inspection of a target power area; wherein the initial historical image set includes multiple initial historical images of the target power area; Adding a defect label to each initial historical image in the initial historical image set based on a preset defect standard to obtain an intermediate historical image set; wherein the intermediate historical image set includes a plurality of labeled intermediate historical images, and the defect label includes a defect location, a defect type, and a defect level; Based on the data enhancement technology, linear interpolation is performed on a plurality of labeled intermediate historical images in the intermediate historical image set, and the historical image set is determined based on the linear interpolation result and the intermediate historical image set.

3. The method according to claim 2, characterized in that The training set includes a training data set and a training label set; and the training of the YOLOv8 defect detection model to be trained based on the training set includes: Obtain the YOLOv8 defect detection model to be trained; wherein the YOLOv8 defect detection model to be trained includes a feature extraction module, a feature fusion module, and a defect prediction module; For each training data in the training data set, performing feature extraction on the training data based on the feature extraction module to obtain feature information of different scales corresponding to the training data; and performing feature fusion processing on the feature information of different scales based on the feature fusion module to obtain feature fusion information; and performing prediction on the feature fusion information based on the defect prediction module to obtain predicted defect information corresponding to the training data; Calculating a loss value between the predicted defect information corresponding to the training data and the defect label of the training sample data corresponding to the training data in the training label set based on a distance loss function, and adjusting the YOLOv8 segmentation model to be trained based on the loss value; Repeat the above steps until the training results meet the preset requirements to obtain the trained YOLOv8 segmentation model.

4. The method according to claim 3, characterized in that The feature extraction module includes a global feature extraction module, a RepResNet feature extraction module, and a feature initial fusion module. The global feature extraction module includes a dilated convolution module and a normal convolution module. The feature extraction of the training data based on the feature extraction module includes: performing feature extraction processing on the training data based on the dilated convolution module to obtain first feature information; performing feature extraction processing on the training data based on the normal convolution module to obtain second feature information; and determining a first feature map based on the first feature information and the second feature information; and performing multi-scale feature enhancement processing on the first feature map based on the RepResNet feature extraction module to determine a first feature representation; and determining first-scale feature information based on the feature initial fusion module, the first feature map, and the first feature representation; performing feature extraction processing on the first-scale feature information based on the dilated convolution module to obtain third feature information; performing feature extraction processing on the first-scale feature information based on the normal convolution module to obtain fourth feature information; and determining a second feature map based on the third feature information and the fourth feature information; and performing multi-scale feature enhancement processing on the second feature map based on the RepResNet feature extraction module to determine a second feature representation; and determining second-scale feature information based on the feature initial fusion module, the second feature map, and the second feature representation; performing feature extraction processing on the second-scale feature information based on the dilated convolution module to obtain fifth feature information; performing feature extraction processing on the second-scale feature information based on the normal convolution module to obtain sixth feature information; and determining a third feature map based on the fifth feature information and the sixth feature information; and performing multi-scale feature enhancement processing on the third feature map based on the RepResNet feature extraction module to determine a third feature representation; and determining third-scale feature information based on the feature initial fusion module, the third feature map, and the third feature representation; Based on the dilated convolution module, feature extraction processing is performed on the third-scale feature information to obtain seventh feature information; and based on the ordinary convolution module, feature extraction processing is performed on the third-scale feature information to obtain eighth feature information; and, based on the seventh feature information and the eighth feature information, a fourth feature map is determined; and, based on the RepResNet feature extraction module, multi-scale feature enhancement processing is performed on the fourth feature map to determine a fourth feature representation; and fourth-scale feature information is determined based on the feature initial fusion module, the fourth feature map, and the fourth feature representation.

5. The method according to claim 4, characterized in that The feature fusion module performs feature fusion processing on the feature information of different scales, including: Based on the feature fusion module, feature fusion processing is performed on the second-scale feature information, the third-scale feature information, and the fourth-scale feature information to obtain the feature fusion information.

6. The method according to claim 3, characterized in that The distance loss function includes: Where L represents the loss value; IoU represents the degree of overlap between the location box of the predicted defect and the location box of the defect label; Distance represents the Euclidean distance measure between the center point of the location box of the predicted defect and the center point of the location box of the defect label.

7. The method according to any one of claims 1 to 6, characterized in that After determining the defect information about the target power area based on the trained YOLOv8 defect detection model and the real-time image data, the method further includes: Determine defect causes and repair suggestions based on a preset defect knowledge base, the defect information, and the real-time image data.

8. A drone inspection device for online detection of power systems, characterized in that: include: An image acquisition module is used to acquire a historical image set of the UAV during inspection of a target power area; wherein the historical image set includes a plurality of labeled historical images of the target power area; A model training module is configured to determine a training set based on the historical image set; and train a YOLOv8 defect detection model to be trained based on the training set to obtain a trained YOLOv8 defect detection model; The defect determination module is used to obtain real-time image data of the drone when inspecting the target power area, and determine defect information about the target power area based on the trained YOLOv8 defect detection model and the real-time image data.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.