Defect detection method, apparatus, device, and medium
By using UAV multi-view acquisition and 3D reconstruction rendering technology, the problems of limited viewpoint and insufficient automation in UAV inspection have been solved, achieving efficient and accurate defect detection and improving the overall efficiency and accuracy of power line inspection.
Patent Information
- Application Number
- CN202511282546.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing drone defect detection technologies suffer from limitations in perspective, low detection efficiency, low automation, and a lack of proactive verification capabilities, resulting in insufficient detection accuracy and efficiency.
By acquiring images from multiple perspectives using drones, preliminary detection is performed. When the confidence level is insufficient, metacognitive processing is triggered to perform 3D reconstruction and rendering, generating a sampling perspective sequence, and then performing image rendering and re-detection to achieve high-precision defect identification.
It improves the accuracy and recall rate of defect detection, shortens operation time, expands the inspection coverage, reduces operating costs, and achieves the best balance between detection efficiency and accuracy.
Smart Images

Figure CN120765656B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power grids, artificial intelligence, etc., and in particular to a defect detection method, device, equipment and medium. BACKGROUND
[0002] The distribution network is an important part of the power system, and its safe and stable operation is crucial. Some key components on the distribution tower, such as insulators and strain clamps, bear the functions of electrical insulation and mechanical support. Due to long-term exposure to complex natural environments, these components may develop cracks, damage, self-explosion, and other defects due to lightning strikes, material aging, external damage, and other factors. If these defects are not discovered and addressed in a timely manner, it may lead to a decrease in line safety and threaten the safety of the power grid. Therefore, how to improve the accuracy and efficiency of defect detection has become a focus of attention.
[0003] The related technology detects defects based on a UAV, and its accuracy is highly dependent on the image quality and viewing angle of a single shot. However, the visibility of component defects on the tower has strong directionality, which causes limitations of a single viewing angle. Another related technology has the UAV perform complex circling flight or multi-point hovering shooting on each tower to be inspected, which greatly increases the inspection operation time of a single tower and significantly consumes the valuable power of the UAV, thereby reducing the overall inspection mileage and efficiency of a single flight. Another related technology performs inspection operations by predefining several viewing angles, which results in a large viewing angle gap between the pictures taken, directly leading to a decrease in detection performance due to factors such as angle and lighting during subsequent detection. Another related technology addresses the viewing angle problem by optimizing hardware and flight strategies, which requires human intervention by a pilot or background personnel to determine the best observation angle, has low automation, and mainly relies on a deep learning-based target detection model to analyze single or multiple frames of images taken by the UAV, lacks picture-to-picture relevance, and lacks a mechanism for active verification when encountering uncertain situations. SUMMARY
[0004] The embodiments of the present application aim to at least solve one of the technical problems in the related art. To this end, one object of the present application is to provide a defect detection method, device, equipment and medium that can improve the accuracy and efficiency of defect detection and actively verify the defect detection results.
[0005] The embodiment of the application provides a defect detection method, which comprises: acquiring an original image of a target object based on a multi-view collection of the target object by a UAV; performing defect detection on the target object based on the original image to obtain a preliminary detection result; in a case where a confidence level of the preliminary detection result indicating that the target object has a defect is less than a preset threshold, triggering a meta-cognition processing mode, performing three-dimensional reconstruction based on the original image to obtain a three-dimensional model of the target object and acquisition parameter information of the original image; obtaining a sampling view sequence of the target object based on the three-dimensional model, the original image and the acquisition parameter information; performing rendering on a three-dimensional image corresponding to the original image in the three-dimensional model based on the sampling view sequence to obtain an image rendering result; and performing defect detection on the target object based on the image rendering result to obtain a defect detection result.
[0006] Exemplarily, the original image is located in an image group; the sampling view sequence of the target object is obtained based on the three-dimensional model, the original image and the acquisition parameter information, which comprises: determining the original image from the image group based on the preliminary detection result, and performing image segmentation on the original image to obtain mask information of the original image; and obtaining the sampling view sequence of the target object based on the three-dimensional model, the mask information of the original image and the acquisition parameter information.
[0007] Exemplarily, the preliminary detection result comprises image index information, first defect confidence level information and position information of the target object in the original image; the original image is determined from the image group based on the preliminary detection result, and image segmentation is performed on the original image to obtain mask information of the original image, which comprises: determining the original image from the image group based on the image index information; and performing image segmentation on the original image based on the position information to obtain the mask information of the original image.
[0008] Exemplarily, the acquisition parameter information of the original image comprises intrinsic information, extrinsic information and depth information, and the sampling view sequence comprises target intrinsic information and target extrinsic information; the sampling view sequence of the target object is obtained based on the three-dimensional model, the mask information of the original image and the acquisition parameter information, which comprises: determining the intrinsic information of the original image as the target intrinsic information; projecting pixel points of the mask information of the original image to a three-dimensional space according to the three-dimensional model based on the intrinsic information, the extrinsic information and the depth information of the original image to obtain three-dimensional point cloud data; obtaining reference extrinsic information based on a center point of the three-dimensional point cloud data; and obtaining the target extrinsic information based on geometric transformation and combination of the reference extrinsic information.
[0009] Exemplarily, the geometric transformation includes depth translation, spherical interpolation sampling, and tangent plane translation; the target extrinsic parameter information is obtained based on the geometric transformation and combination based on the reference extrinsic parameter information, including: performing depth translation along the depth axis direction with the reference extrinsic parameter information as the center to obtain a first extrinsic parameter; performing coordinate interpolation movement on the sphere with the reference extrinsic parameter information as the sphere center to obtain a second extrinsic parameter; performing tangent plane translation along the direction perpendicular to the central plane with the reference extrinsic parameter information as the center to obtain a third extrinsic parameter; and obtaining the target extrinsic parameter information based on the combination between any one or more of the first extrinsic parameter, the second extrinsic parameter, and the third extrinsic parameter.
[0010] Exemplarily, based on the sampling view angle sequence, the three-dimensional image corresponding to the original image in the three-dimensional model is rendered to obtain an image rendering result, including: taking the center point of the three-dimensional point cloud data as the rendering center, rendering the three-dimensional image corresponding to the original image in the three-dimensional model to obtain the image rendering result.
[0011] Exemplarily, the image rendering result includes image index information corresponding to the preliminary detection result, first defect confidence information, and position information; based on the image rendering result, the target object is detected to obtain a defect detection result, including: based on the image rendering result, the target object is detected to obtain second defect confidence information; in the case that the second defect confidence information indicates that the confidence of the target object existing defects is less than a preset threshold, the corresponding preliminary detection result is removed based on the image index information to obtain the defect detection result.
[0012] Exemplarily, the defect detection result includes image index information corresponding to the preliminary detection result, first defect confidence information corresponding to the preliminary detection result or second defect confidence information corresponding to the defect detection result, and position information corresponding to the preliminary detection result; the method further includes: based on the image index information, the defect detection result is structured and organized to obtain a defect report; and the defect report is displayed.
[0013] Exemplarily, the target object includes a power distribution tower; the original image is obtained by multi-view collection of the target object based on a drone, including: collecting front view, left view, right view, top view of the power distribution tower, and overall view of the power distribution tower based on the drone to obtain the original image.
[0014] Another embodiment of the present application provides a defect detection device, comprising: an acquisition module configured to acquire original images of a target object based on multi-view acquisition by a UAV; a first detection module configured to detect defects of the target object based on the original images to obtain preliminary detection results; a reconstruction module configured to trigger metacognition processing mode when a confidence level of the preliminary detection results indicating that the target object has defects is less than a preset threshold, and perform three-dimensional reconstruction based on the original images to obtain a three-dimensional model of the target object and acquisition parameter information of the original images; an obtaining module configured to obtain a sampling view sequence of the target object based on the three-dimensional model, the original images and the acquisition parameter information; a rendering module configured to render a three-dimensional image corresponding to the original images in the three-dimensional model based on the sampling view sequence to obtain image rendering results; and a second detection module configured to detect defects of the target object based on the image rendering results to obtain defect detection results.
[0015] Another embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method of any of the above embodiments when executing the computer program.
[0016] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method of any of the above embodiments.
[0017] In the above embodiments, the defect detection method comprises: acquiring original images of a target object based on multi-view acquisition by a UAV; detecting defects of the target object based on the original images to obtain preliminary detection results; triggering metacognition processing mode when a confidence level of the preliminary detection results indicating that the target object has defects is less than a preset threshold, and performing three-dimensional reconstruction based on the original images to obtain a three-dimensional model of the target object and acquisition parameter information of the original images; obtaining a sampling view sequence of the target object based on the three-dimensional model, the original images and the acquisition parameter information; rendering a three-dimensional image corresponding to the original images in the three-dimensional model based on the sampling view sequence to obtain image rendering results; and detecting defects of the target object based on the image rendering results to obtain defect detection results. The above triggers metacognition processing when the confidence level of the defects is less than the preset threshold, and performs three-dimensional reconstruction to obtain the sampling view sequence, actively renders and re-detects from the sampling view, which can effectively "see through" the occlusion and "get around" the unfavorable reflection, thereby accurately identifying hidden and small defects (such as side cracks and hidden explosions) that are difficult to find under the original shooting view, greatly improving the accuracy and recall rate of defect detection; and significantly shortening the operation time of a single foundation tower, improving the inspection coverage range of a single flight, directly reducing the operation cost of power inspection, and achieving the best balance between detection efficiency and accuracy.
[0018] Additional aspects and advantages of the present application will be apparent from the following description, which illustrates by way of example preferred embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 A defect detection method flowchart provided for an embodiment of the present application;
[0020] Figure 2 A defect detection method flowchart provided for an embodiment of the present application;
[0021] Figure 3 A defect detection device block diagram provided for another embodiment of the present application;
[0022] Figure 4 An electronic device block diagram provided for another embodiment of the present application. DETAILED DESCRIPTION
[0023] Embodiments of the present application are described in detail below with reference to the attached drawings, which show by way of example embodiments in which the same or similar elements or elements having the same or similar functions are denoted by the same or similar reference numerals throughout the drawings. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.
[0024] The power distribution network is an important part of the power system, and its safe and stable operation is crucial. Some key components on the distribution tower, such as insulators, strain clamps, etc., bear the functions of electrical insulation, mechanical support, etc. Due to long-term exposure to complex natural environment, these components may develop cracks, damage, self-explosion, etc. due to lightning strikes, material aging, external force damage, etc. If these defects are not discovered and handled in time, it may lead to a decrease in line safety and threaten the safety of the power grid.
[0025] In recent years, using unmanned aerial vehicles (UAVs) to carry high-definition cameras for power transmission and distribution line inspection has become a mainstream trend. Compared with traditional manual tower climbing operations, UAV inspection has the advantages of high efficiency, low cost, good safety, and no terrain restrictions. During the inspection process, the UAV flies along the predetermined route, takes pictures of the tower and key components, and then the intelligent recognition algorithm in the background automatically detects defects.
[0026] However, the accuracy of related UAV vision-based detection methods is highly dependent on the quality and angle of the single-shot image, while the visibility of component defects on the tower has a strong directionality. For example, if a piece of insulator explodes (the umbrella skirt is damaged), if the shooting angle of the UAV is directly opposite or opposite to the damaged surface, the defect is clearly visible; but if it is shot from the side (i.e. along the radial direction of the insulator), the damage feature may be blocked by the intact umbrella skirt structure.
[0027] To circumvent the limitation of single perspective, current industry practice usually requires the UAV to perform a complex fly-around or multi-point hovering shooting for each tower to be inspected to obtain as many observation angles as possible. Although this approach can improve the detection rate to some extent, it greatly increases the inspection time of single-base towers, significantly consumes the valuable power of the UAV, and thus reduces the overall inspection mileage and efficiency of a single flight. In addition, there is a part of the inspection operation by predefining several perspectives (usually 5, which are front view, left view, right view, tower top (i.e. top view), and tower overall view). Although such shooting is faster, the large difference in perspective between the pictures will directly lead to a decrease in detection performance due to factors such as angle and lighting. Therefore, how to improve the efficiency of UAV inspection while ensuring high detection rate and high accuracy is a technical pain point that needs to be solved in the field.
[0028] Related technologies mainly deal with the perspective problem by optimizing hardware and flight strategies. For example, a UAV equipped with a high-magnification zoom gimbal is used to zoom in on suspicious points for closer observation. However, this still requires human intervention by the pilot or background personnel to determine the best observation angle, and the degree of automation is not high. At the algorithm level, the main approach is to rely on deep learning-based target detection models to analyze single or multiple frames of images taken by the UAV, which lacks correlation between pictures and lacks a mechanism for active verification in uncertain situations. The problems are as follows: (1) The contradiction between inspection efficiency and detection accuracy: To ensure high accuracy and low missed detection rate, the UAV needs to perform time-consuming and power-consuming fly-around, resulting in low inspection efficiency. If a fast fly-by approach is used, the detection accuracy will be sacrificed due to the single shooting angle. (2) Limited intelligence level, lack of active verification ability: existing algorithms are passive in reasoning. The model cannot actively seek more information to support or refute its initial judgment for a target that "looks like but is uncertain", i.e. lacks a kind of "meta-cognition" reflection ability similar to that of human experts.
[0029] Therefore, the embodiments of the present application provide a defect detection method. Instead of treating the images taken by the UAV at different points (referred to as original shooting images) as independent detection tasks, the images are input as a sparse, multi-perspective scene to construct a high-precision local three-dimensional digital twin. When preliminary analysis of the original shooting images indicates the presence of uncertain and suspicious defects, the system will trigger a meta-cognition process: using the constructed three-dimensional model, the best sampling perspective is found based on the current uncertain defect, and then new pictures are rendered for detection. Through this "virtual close-up detailed inspection", uncertainty is eliminated, and a multi-perspective verified and highly reliable detection conclusion (defect detection result) is formed.
[0030] Figure 1A defect detection method flowchart is provided for an embodiment of the present application.
[0031] As shown in Figure 1 , the defect detection method 100 includes steps S110-S160.
[0032] Step S110, based on the unmanned aerial vehicle, multi-view collection is performed on the target object to obtain an original image.
[0033] Exemplarily, the target object includes a power distribution tower, for example, and the multi-view collection can be an elevation view, a left view, a right view, a power distribution tower top, and a power distribution tower overall view, etc. of the power distribution tower. The original image is an image sequence I composed of multi-view images, containing multiple images, such as a picture sequence I={ } consisting of N pictures.
[0034] Step S120, based on the original image, the target object is subjected to defect detection to obtain a preliminary detection result.
[0035] Exemplarily, the defect detection uses a defect detection model M1, which can be a trained yolov5 component defect detection model M1. The specific architecture adopted by the defect detection model M1 is not specifically limited. The preliminary detection result R can include picture index information (for corresponding the preliminary detection result R with the original image), confidence information (for indicating the reliability of the preliminary detection result R), and position information (for indicating the specific position of the preliminary detection result R). For the input picture sequence I, its corresponding preliminary detection result R can include multiple, such as the preliminary detection result R={ }, wherein R1 represents the detection result of I1 in the original image sequence, R2 represents the detection result of I2 in the original image sequence, etc. represents the picture index (i.e. which original picture it corresponds to), represents the confidence of M1 determining that the target belongs to a defect, represents the position of M1 determining the target.
[0036] Step S130, in the case where the preliminary detection result indicates that the confidence of the target object existing a defect is less than a preset threshold, a meta-cognition processing mode is triggered, and based on the original image, three-dimensional reconstruction is performed to obtain a three-dimensional model of the target object and acquisition parameter information of the original image.
[0037] For example, the acquisition parameter information of the original image may include intrinsic parameter information, extrinsic parameter information, and depth information. The preliminary detection results include a certain result and a result to be verified (distinguished based on the confidence information in the preliminary detection results). When the defect confidence information of the preliminary detection result is higher than a preset threshold T (certain result), the preliminary detection result is taken as the final defect detection result, and the certain result is marked as r_certain={ ,..., When the confidence level of the defects in the preliminary detection results is lower than the preset threshold T (result to be verified), the preliminary detection results need to be further confirmed. The result to be verified is denoted as r_uncertain=
[0038] { ,... The algorithm will actively trigger metacognitive processing, obtaining the original image corresponding to the preliminary detection result based on the index information. This original image is then input into the forward 3D reconstruction algorithm to obtain the reconstructed 3D model M2 of the power distribution tower, as well as the camera intrinsic parameters K, extrinsic parameters P, and depth map D corresponding to the image in I. The intrinsic parameter information is shared, and the extrinsic parameter information P = { } and depth information D={ Each original image has a publicly available 3D reconstruction algorithm, which can be a method such as AnySplat that supports 3D scene reconstruction from sparse multi-view input while predicting camera intrinsics, extrinsics, and depth maps.
[0039] Step S140: Based on the 3D model, the original image, and the acquisition parameter information, obtain the sampling view sequence of the target object.
[0040] For example, the sampling view sequence includes intrinsic and extrinsic information. Based on the original image corresponding to the preliminary detection results, the mask of the corresponding target is obtained using the publicly available Segment All (SAM) model. The reference extrinsic information can be obtained by back-projecting the pixels on the mask based on the acquired parameter information. The sampling view sequence can be obtained by performing geometric transformation on the 3D model based on the reference extrinsic information. The sampling view sequence may include multiple samples.
[0041] Step S150: Based on the sampling view sequence, render the 3D image corresponding to the original image in the 3D model to obtain the image rendering result.
[0042] For example, based on the obtained multiple sampling view sequences, the three-dimensional image corresponding to the original image in the three-dimensional model is transformed into a high-quality two-dimensional image under a specific view (sampling view sequence), and the image rendering result includes two-dimensional images corresponding to multiple sampling view sequences.
[0043] At step S160, based on the image rendering result, the target object is subjected to defect detection to obtain a defect detection result.
[0044] Exemplarily, the defect detection result includes picture index information, confidence information and position information, the image rendering result is subjected to defect detection using a defect detection model M1, which can be a trained yolov5 component defect detection model M1, the specific architecture adopted by the defect detection model M1 is not specifically limited, the picture index information (used for corresponding the result obtained using the defect detection model with the original image), the confidence information (used for indicating the reliability of the result obtained using the defect detection model) and the position information (used for indicating the specific position of the result obtained using the defect detection model) are obtained, based on the confidence information and a preset threshold T, the to-be-verified result corresponding to less than the preset threshold is eliminated, and the result higher than the preset threshold T in the to-be-verified result and the confident result higher than the preset threshold T in step S130 are taken as the final defect detection result.
[0045] In the above embodiment, when the confidence of the defect is less than the preset threshold, metacognition processing is triggered, and three-dimensional reconstruction is performed to obtain a sampling view sequence, and active rendering and re-detection are performed from the sampling view, which can effectively "see through" the occlusion and "get around" the unfavorable reflection, thereby accurately identifying hidden and small defects (such as side cracks and hidden explosions) that are difficult to find under the original shooting view, greatly improving the accuracy and recall rate of defect detection; and significantly shortening the operation time of a single foundation tower, improving the inspection coverage range of a single flight, directly reducing the operation cost of power inspection, starting the three-dimensional analysis and virtual rendering process with high cost only for the low-confidence target to be verified, and achieving the best balance between detection efficiency and accuracy.
[0046] The original image is located in an image group; based on the three-dimensional model, the original image and the acquisition parameter information, a sampling view sequence of the target object is obtained, including: indexing and determining the original image from the image group based on the preliminary detection result, and performing image segmentation on the original image to obtain mask information of the original image; based on the three-dimensional model, the mask information of the original image and the acquisition parameter information, the sampling view sequence of the target object is obtained.
[0047] Specifically, the image segmentation adopts a public segmentation all model (SAM) to obtain a mask corresponding to the target, based on the acquisition parameter information, the pixel points on the mask are subjected to back projection to obtain reference external parameter information, based on the reference external parameter information, the three-dimensional model is subjected to geometric transformation to obtain the sampling view sequence, wherein the sampling view sequence can include multiple.
[0048] Exemplarily, the preliminary detection result includes image index information, first defect confidence information, and position information of the target object in the original image; the original image is indexed and determined from the image group based on the preliminary detection result, and image segmentation is performed on the original image to obtain mask information of the original image, including: indexing and determining the original image from the image group based on the image index information; performing image segmentation on the original image based on the position information to obtain mask information of the original image.
[0049] For example, for each element in r_uncertain , according to , the original image is obtained , the image and (position information) are input into the SAM to obtain the mask information of the corresponding target.
[0050] Exemplarily, the acquisition parameter information of the original image includes intrinsic information, extrinsic information, and depth information, and the sampling view sequence includes target intrinsic information and target extrinsic information; based on the three-dimensional model, the mask information of the original image, and the acquisition parameter information, the sampling view sequence of the target object is obtained, including: determining the intrinsic information of the original image as the target intrinsic information; based on the intrinsic information, the extrinsic information, and the depth information of the original image, the pixel points of the mask information of the original image are back projected to the three-dimensional space according to the three-dimensional model to obtain three-dimensional point cloud data; based on the center point of the three-dimensional point cloud data, reference extrinsic information is obtained; based on the reference extrinsic information, geometric transformation and combination are performed to obtain target extrinsic information.
[0051] Specifically, the intrinsic information of the original image is shared, all original images correspond to the same intrinsic information, the geometric transformation can include depth translation, spherical interpolation sampling, and tangent plane translation, and the target extrinsic information includes multiple target extrinsic information .
[0052] For example, using the camera pose {K, } and the depth information , the pixel points on the two-dimensional mask mask_2d are back projected (unprojected) to the three-dimensional space to obtain a set of three-dimensional point clouds describing the spatial position of the defect; the center point of is calculated, and the center point is set as the position of the center point of the rendered image, and then the camera extrinsic information (reference extrinsic information) facing the center of the defect can be obtained; according to , a set of preset geometric transformation strategies (depth translation, spherical interpolation sampling, and tangent plane translation) are used to generate sampling views with diversity (sampling view sequence).
[0053] Exemplarily, the geometric transformations include depth translation, spherical interpolation sampling, and tangent plane translation; the target extrinsic parameter information is obtained based on the reference extrinsic parameter information and the geometric transformations and combinations, including: performing depth translation along the depth axis direction with the reference extrinsic parameter information as the center to obtain a first extrinsic parameter; performing coordinate interpolation movement on a spherical surface with the reference extrinsic parameter information as the spherical center to obtain a second extrinsic parameter; performing tangent plane translation along a direction perpendicular to the center plane with the reference extrinsic parameter information as the center to obtain a third extrinsic parameter; and obtaining the target extrinsic parameter information based on combinations between any one or more of the first extrinsic parameter, the second extrinsic parameter, and the third extrinsic parameter.
[0054] Specifically, the target extrinsic parameter information can be the first extrinsic parameter, the second extrinsic parameter, or the third extrinsic parameter, or a combination of the first extrinsic parameter and the second extrinsic parameter, a combination of the second extrinsic parameter and the third extrinsic parameter, or a combination of the second extrinsic parameter and the third extrinsic parameter, or a combination of the first extrinsic parameter, the second extrinsic parameter, and the third extrinsic parameter, wherein the first extrinsic parameter is a plurality of extrinsic parameters obtained by moving with the reference extrinsic parameter information as the center, the second extrinsic parameter is a plurality of extrinsic parameters obtained by moving with the reference extrinsic parameter information as the spherical center, and the third extrinsic parameter is a plurality of extrinsic parameters obtained by translating along a direction perpendicular to the tangent plane with the reference extrinsic parameter information as the center. The geometric transformations are not limited to depth translation, spherical interpolation sampling, and tangent plane translation, and can also include other geometric transformation forms.
[0055] For example, according to the reference extrinsic parameter information, a set of preset geometric transformation strategies are used to generate sampling view angles with diversity, which include but are not limited to:
[0056] a) Depth translation: moving the camera along its main optical axis (i.e., the direction towards the defect center) by a certain distance to simulate the observation effect of “zooming in” and “zooming out”.
[0057] b) Spherical interpolation sampling: moving the camera position on a virtual spherical surface with the defect center point as the spherical center by spherical coordinate interpolation, while keeping the camera always facing the spherical center to simulate the effect of surrounding observation.
[0058] c) Tangent plane translation: moving the camera position on a plane perpendicular to its main optical axis up, down, left, and right to observe the morphology of the defect under different light incidence angles.
[0059] By combining these strategies, ultimately, u_A new camera extrinsic parameters are obtained, which constitute the complete sampling view angle set of the defect to be verified (target extrinsic parameter information) and (target extrinsic parameter information), where { }, The intrinsic parameter information K of the direct original image is used to obtain A sets of sampling perspectives (an intrinsic parameter and an extrinsic parameter together determine a shooting angle and field of view, i.e., the shooting scene is determined).
[0060] In the above embodiments, 2D segmentation masks, camera poses and 3D model geometric information are used to jointly determine one or a set of virtual observation perspectives that can maximize the defect information gain, which greatly improves the accuracy and recall of defect detection.
[0061] Based on the sampling view sequence, the three-dimensional image corresponding to the original image in the three-dimensional model is rendered to obtain the image rendering result, including: using the center point of the three-dimensional point cloud data as the rendering center, the three-dimensional image corresponding to the original image in the three-dimensional model is rendered to obtain the image rendering result.
[0062] Specifically, based on the obtained multiple sampling view sequences, the 3D images corresponding to the original images in the 3D model are... It is transformed into a high-quality two-dimensional image under a specific viewpoint (sampling viewpoint sequence). The image rendering result includes two-dimensional images corresponding to multiple sampling viewpoint sequences.
[0063] For example, with Centered on the rendering center, based on the 3D model M2 and the sampling view sequence , Render the corresponding image sequence This yields the image rendering result.
[0064] The image rendering result includes image index information, first defect confidence information, and location information corresponding to the preliminary detection result; based on the image rendering result, defect detection is performed on the target object to obtain the defect detection result, including: based on the image rendering result, defect detection is performed on the target object to obtain second defect confidence information; if the confidence of the second defect confidence information indicating that the target object has a defect is less than a preset threshold, the corresponding preliminary detection result is removed based on the image index information to obtain the defect detection result.
[0065] Specifically, the first defect confidence information includes confidence information corresponding to the preliminary detection result, the defect detection is performed using a defect detection model M1, the defect detection model can be a trained yolov5 component defect detection model M1, and the specific architecture of the defect detection model M1 is not limited, and the picture index information (used to correspond the result obtained by using the defect detection model to the original image), the confidence information (used to represent the reliability of the result obtained by using the defect detection model) and the position information (used to represent the specific position of the result obtained by using the defect detection model) are obtained. Based on the confidence information and the preset threshold T, the to-be-verified result less than the preset threshold is removed, and the result higher than the preset threshold T in the to-be-verified result and the certain result higher than the preset threshold T in step S130 are used as the final defect detection result.
[0066] For example, the rendered image is detected using the M1 model, and the corresponding detection result is output, the highest confidence is obtained, if the highest confidence is still lower than the preset threshold T, it is considered that there is no defect, and the Finally, the target set r_uncertain_update (removed to-be-verified result) removing all non-defects is obtained. The “certain” detection result r_certain and the result r_uncertain_update after the metacognition verification process are integrated, and the final result is output.
[0067] The defect detection result includes image index information corresponding to the preliminary detection result, first defect confidence information corresponding to the preliminary detection result or second defect confidence information corresponding to the defect detection result, and position information corresponding to the preliminary detection result. The method further comprises: based on the image index information, the defect detection result is structurally organized and processed to obtain a defect report; and the defect report is displayed.
[0068] For example, the “certain” defect set r_certain and the verified defect set r_uncertain_update are integrated to form a final, high-reliability defect report. The integration method is as follows: first, all defect records in r_certain are directly included in the final report. Second, each verified and confirmed defect record in r_uncertain_update is also included in the final report. The final report R_new can be structurally organized according to the original picture index img_index, for example, for each original picture, all “certain” defects and verified and confirmed defects contained thereon are listed, and the final, high-confidence score and accurate position are attached.
[0069] In the above embodiment, by effectively integrating the "confirmed" results from the original image and the "verified" results from the virtual image, a final high-reliability report is formed, and the defect detection result can be more intuitively observed.
[0070] The target object includes a power distribution pole tower; the original image is obtained by multi-view collection based on the unmanned aerial vehicle, including: collecting the front view, left view, right view, top view and overall view of the power distribution pole tower based on the unmanned aerial vehicle to obtain the original image.
[0071] Exemplarily, the multi-view collection can be the front view, left view, right view, top view and overall view of the power distribution pole tower, and the like, and the original image is an image sequence I composed of multi-view images, containing multiple images, such as a picture sequence I={ } composed of N pictures.
[0072] In the above embodiment, by constructing a high-fidelity three-dimensional digital twin and actively rendering and re-detecting from the "best" virtual perspective, the occlusion can be effectively "seen through" and the unfavorable reflection can be "circled around", so that hidden and small defects (such as side cracks and hidden self-explosions) that are difficult to find in the original shooting perspective can be accurately identified, and the accuracy and recall rate of defect detection are greatly improved; and through the pure algorithm means of "virtual close-up detailed inspection", multi-angle analysis of suspicious points can be realized on the server side. This makes the unmanned aerial vehicle only need to perform efficient and fast standard route shooting in the front end, without the need to repeatedly adjust the posture and position for a single suspicious point, thereby significantly shortening the operation time of a single base pole tower, improving the inspection coverage range of a single flight, directly reducing the operation cost of power inspection, and through the "meta-cognition" working paradigm, the "where to look" (best perspective) and "how to look" (geometric transformation strategy) can be autonomously planned, which is an intelligent upgrade from passive response to active exploration, making the detection system closer to the thinking mode of human experts.
[0073] Figure 2 A defect detection method provided for an embodiment of the present application is shown in a whole flowchart.
[0074] As shown in Figure 2 , the defect detection method 200 includes steps S201-S208.
[0075] Step S201, picture sequence photographed at different points of the same tower pole.
[0076] Exemplarily, for the image group I photographed for the same tower, the image group I of one tower in the embodiment is composed of five predetermined perspectives (front view, left view, right view, tower top, and tower overall view), and the composition of the image group can also be other ways.
[0077] Step S202, the defect detection model M1 obtains a detection result R.
[0078] Exemplarily, the images in the image group I are input into the first-order component defect detection model M1 one by one to obtain a preliminary detection result R.
[0079] Step S203A, a to-be-verified set r_certain is obtained according to the threshold.
[0080] Step S203B, a to-be-verified set r_uncertain is obtained according to the threshold.
[0081] Exemplarily, the detection result is analyzed, the result with a confidence lower than a preset threshold T is marked as “to be verified” and subsequent steps S204-S207 are performed, and the rest of the detection result is marked as “certain”.
[0082] Step S204, a three-dimensional reconstruction algorithm obtains corresponding intrinsic parameters, extrinsic parameters, a depth map and a three-dimensional model.
[0083] Step S205, it is judged whether the result in r_uncertain is traversed to an end, if not, it is turned to S206, and if yes, the result is integrated.
[0084] Step S206, a best sampling view sequence is calculated and a picture is rendered, and a detection result is obtained by using M1 detection.
[0085] Step S207, it is judged whether the highest confidence of the detection result is greater than a threshold, if not, r_uncertain is updated, and if yes, it is turned to step S205.
[0086] Step S208, the result is integrated.
[0087] Exemplarily, the “certain” detection result r_certain and the result r_uncertain_update after the metacognition verification process are integrated, and a final result is output.
[0088] In the above embodiment, through the multi-view preliminary detection→uncertainty recognition based on confidence and metacognition triggering→three-dimensional scene reconstruction based on multi-view input→active best virtual view planning based on 2D / 3D information linkage→rendering, re-detection and final decision based on virtual view→result integration output closed loop working process, the best balance of detection efficiency and accuracy is realized.
[0089] Figure 3 A defect detection device block diagram for another embodiment of the application.
[0090] An embodiment of the application provides a defect detection device 300, please refer to Figure 3The defect detection apparatus 300 comprises: an acquisition module 310, a first detection module 320, a reconstruction module 330, an obtaining module 340, a rendering module 350, and a second detection module 360.
[0091] The acquisition module 310 is configured to acquire the original image based on multi-view acquisition of the target object by the UAV.
[0092] The first detection module 320 is configured to perform defect detection on the target object based on the original image to obtain a preliminary detection result.
[0093] The reconstruction module 330 is configured to, in a case where a confidence level of the preliminary detection result indicating that the target object has a defect is less than a preset threshold, trigger a meta-cognition processing mode, perform three-dimensional reconstruction based on the original image to obtain a three-dimensional model of the target object and acquisition parameter information of the original image.
[0094] The obtaining module 340 is configured to obtain a sampling view sequence of the target object based on the three-dimensional model, the original image, and the acquisition parameter information.
[0095] The rendering module 350 is configured to perform rendering on a three-dimensional image corresponding to the original image in the three-dimensional model based on the sampling view sequence to obtain an image rendering result.
[0096] The second detection module 360 is configured to perform defect detection on the target object based on the image rendering result to obtain a defect detection result.
[0097] It can be understood that the specific description of the defect detection apparatus 300 can refer to the description of the defect detection method in the foregoing, and will not be repeated here.
[0098] The original image is located in an image group; the obtaining module 340 is further configured to determine the original image from the image group based on the preliminary detection result, perform image segmentation on the original image to obtain mask information of the original image, and obtain the sampling view sequence of the target object based on the three-dimensional model, the mask information of the original image, and the acquisition parameter information.
[0099] The preliminary detection result comprises image index information, first defect confidence level information, and position information of the target object in the original image; the obtaining module 340 is further configured to determine the original image from the image group based on the image index information, perform image segmentation on the original image based on the position information to obtain mask information of the original image.
[0100] Exemplarily, the acquisition parameter information of the original image includes intrinsic information, extrinsic information and depth information, and the sampling view sequence includes target intrinsic information and target extrinsic information; the obtaining module 340 is further configured to determine the intrinsic information of the original image as the target intrinsic information; based on the intrinsic information, the extrinsic information and the depth information of the original image, the pixel points of the mask information of the original image are back-projected to a three-dimensional space according to the three-dimensional model to obtain three-dimensional point cloud data; based on the center point of the three-dimensional point cloud data, reference extrinsic information is obtained; based on the reference extrinsic information, geometric transformation and combination are performed to obtain the target extrinsic information.
[0101] Exemplarily, the geometric transformation includes depth translation, spherical interpolation sampling and tangent plane translation; the obtaining module 340 is further configured to perform depth translation along a depth axis direction with the reference extrinsic information as the center to obtain a first extrinsic information; perform coordinate interpolation movement on a spherical surface with the reference extrinsic information as the spherical center to obtain a second extrinsic information; perform tangent plane translation along a direction perpendicular to a central plane with the reference extrinsic information as the center to obtain a third extrinsic information; based on combination between any one or more of the first extrinsic information, the second extrinsic information and the third extrinsic information, the target extrinsic information is obtained.
[0102] Exemplarily, the rendering module 350 is further configured to take the center point of the three-dimensional point cloud data as a rendering center to render a three-dimensional image corresponding to the original image in the three-dimensional model to obtain an image rendering result.
[0103] Exemplarily, the image rendering result includes image index information corresponding to the preliminary detection result, first defect confidence information and position information; the second detection module 360 is further configured to perform defect detection on the target object based on the image rendering result to obtain second defect confidence information; in a case where the second defect confidence information indicates that a confidence that the target object has a defect is less than a preset threshold, the corresponding preliminary detection result is removed based on the image index information to obtain a defect detection result.
[0104] Exemplarily, the defect detection result includes image index information corresponding to the preliminary detection result, first defect confidence information corresponding to the preliminary detection result or second defect confidence information corresponding to the defect detection result, and position information corresponding to the preliminary detection result; the defect detection apparatus 300 further includes: based on the image index information, the defect detection result is structured and organized to obtain a defect report; and the defect report is displayed.
[0105] Exemplarily, the target object includes a power distribution pole tower; the acquisition module 310 is further configured to acquire a front view, a left view, a right view, a top view of the power distribution pole tower and an overall view of the power distribution pole tower based on the unmanned aerial vehicle to obtain the original image.
[0106] Figure 4 An electronic device block diagram is provided for another embodiment of the present application.
[0107] The embodiments of the present application provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the above method when executing the computer program.
[0108] As shown in Figure 4 For ease of understanding, the embodiments of the present application show a specific electronic device 400.
[0109] The electronic device 400 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0110] As shown in Figure 4 The electronic device 400 includes a computing unit 401 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded into a random access memory (RAM) 403 from a storage unit 408. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0111] A plurality of components in the electronic device 400 are connected to the input / output (I / O) interface 405, including an input unit 406, such as a keyboard, a mouse, etc., an output unit 407, such as various types of displays, a speaker, etc., a storage unit 408, such as a magnetic disk, an optical disk, etc., and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0112] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs the various methods described above. For example, in some embodiments, any one or more of the various methods described above can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of any one or more of the various methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform any one or more of the various methods described above by other any appropriate means, such as by means of firmware.
[0113] An embodiment of the present application provides a computer readable storage medium, having stored thereon a computer program, which when executed by a processor implements the steps of the method of any of the above embodiments.
[0114] It is to be appreciated that the above description and the examples that follow are intended to be illustrative only and that changes can be made to the description, as represented by the above listed elements, by the steps recited in the flow charts, and by the examples that follow, without departing from the spirit of the application. Accordingly, the scope of the present application is intended to be defined only by the appended claims and equivalents thereof.
[0115] It should be understood that aspects of the application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, can be used: a hybrid of the technologies mentioned above, discrete logic circuitry having logic gates for implementing logic functions upon data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), or the like.
[0116] In the description of the present application, the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" are intended to mean that a particular feature, structure, material, or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. The illustrative appearances of the above terms in various places in the specification are not intended to exclude that the same feature, structure, material, or characteristic can be present not only in one embodiment or example but in all embodiments and examples of the present application. In other words, the described particular feature, structure, material, or characteristic is not necessarily limited to a single embodiment or example.
[0117] In the description of the present application, it needs to be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.
[0118] In addition, the terms "first", "second", and the like used in the embodiments of the present application are only for the purpose of description, and cannot be understood as indicating or implying relative importance, or implicitly indicating the number of technical features referred to in the embodiments. Therefore, the features defined with "first", "second" and the like in the embodiments of the present application can be explicitly or implicitly indicated to include at least one of the features. In the description of the present application, the meaning of the word "plurality" is at least two or two or more, such as two, three, four, etc., unless otherwise specifically limited in the embodiments.
[0119] In the present application, unless otherwise specifically provided or limited in the embodiments, the terms "mounting", "connecting", "connecting" and "fixing" and the like appearing in the embodiments should be understood broadly, for example, the connection can be fixed connection, or detachable connection, or integral, which can be understood, or can be mechanical connection, electrical connection, etc. Of course, it can also be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements, or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific implementation situation.
[0120] In the present application, unless otherwise specifically provided and limited, the first feature "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature "above", "above" and "above" the second feature can be that the first feature is directly above or obliquely above the second feature, or only indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature "below", "below" and "below" the second feature can be that the first feature is directly below or obliquely below the second feature, or only indicates that the horizontal height of the first feature is less than that of the second feature.
[0121] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that variations, modifications, substitutions and changes can be made by those skilled in the art without departing from the scope of the present application.
Claims
1. A defect detection method characterized by, The method includes: Original images are obtained by using drones to capture target objects from multiple perspectives. Based on the original image, defect detection is performed on the target object to obtain preliminary detection results; If the confidence level of the preliminary detection result indicating that the target object has a defect is less than a preset threshold, the metacognitive processing mode is triggered to perform three-dimensional reconstruction based on the original image, thereby obtaining the three-dimensional model of the target object and the acquisition parameter information of the original image; Based on the 3D model, the original image, and the acquisition parameter information, the sampling view sequence of the target object is obtained; Based on the sampling view sequence, the three-dimensional image corresponding to the original image in the three-dimensional model is rendered to obtain the image rendering result; Based on the image rendering results, defect detection is performed on the target object to obtain defect detection results.
2. The defect detection method according to claim 1, characterized by, The original image is located in the image group; The step of obtaining the sampling viewpoint sequence of the target object based on the 3D model, the original image, and the acquisition parameter information includes: Based on the preliminary detection results, the original image is determined by indexing from the image group, and the original image is segmented to obtain the mask information of the original image; Based on the 3D model, the mask information of the original image, and the acquisition parameter information, the sampling view sequence of the target object is obtained.
3. The defect detection method according to claim 2, characterized by, The preliminary detection results include image index information, first defect confidence information, and the location information of the target object in the original image; Based on the preliminary detection results, the original image is indexed from the image group and determined. Image segmentation is then performed on the original image to obtain its mask information, including: Based on the image index information, the original image is determined from the image group by indexing; Based on the location information, the original image is segmented to obtain the mask information of the original image.
4. The method of claim 2, wherein, The acquisition parameter information of the original image includes intrinsic parameter information, extrinsic parameter information and depth information, and the sampling viewpoint sequence includes target intrinsic parameter information and target extrinsic parameter information; The step of obtaining the sampling viewpoint sequence of the target object based on the 3D model, the mask information of the original image, and the acquisition parameter information includes: The intrinsic parameter information of the original image is used to determine the target intrinsic parameter information; Based on the intrinsic parameter information, extrinsic parameter information, and depth information of the original image, the pixel points of the mask information of the original image are back-projected into the three-dimensional space according to the three-dimensional model to obtain three-dimensional point cloud data; Based on the center point of the three-dimensional point cloud data, reference extrinsic information is obtained; Geometric transformations and combinations are performed based on the reference extrinsic information to obtain the target extrinsic information.
5. The defect detection method of claim 4, wherein The geometric transformation includes depth translation, spherical interpolation sampling, and tangent plane translation; the geometric transformation and combination based on the reference extrinsic information to obtain the target extrinsic information includes: Using the reference extrinsic information as the center, perform a depth translation along the depth axis to obtain the first extrinsic parameter; Using the reference extrinsic information as the center of the sphere, coordinate interpolation is performed on the sphere to obtain the second extrinsic parameter; The third external parameter is obtained by cutting and translating along a direction perpendicular to the central plane with the reference external parameter information as the center; The target external parameter information is obtained based on a combination between any one or more of the first external parameter, the second external parameter, and the third external parameter.
6. The defect detection method of claim 4, wherein The image rendering result is obtained by rendering a three-dimensional image corresponding to the original image in the three-dimensional model based on the sampling view sequence, including: The image rendering result is obtained by rendering a three-dimensional image corresponding to the original image in the three-dimensional model with a center point of the three-dimensional point cloud data as a rendering center.
7. The defect detection method according to claim 3, characterized by, The image rendering result includes image index information, first defect confidence information, and position information corresponding to the preliminary detection result; The defect detection result is obtained by performing defect detection on the target object based on the image rendering result, including: The second defect confidence information is obtained by performing defect detection on the target object based on the image rendering result; In a case where the second defect confidence information indicates that a confidence that the target object has a defect is less than the preset threshold, the defect detection result is obtained by removing the corresponding preliminary detection result based on the image index information.
8. The defect detection method of claim 7, wherein The defect detection result includes image index information corresponding to the preliminary detection result, first defect confidence information corresponding to the preliminary detection result or second defect confidence information corresponding to the defect detection result, and position information corresponding to the preliminary detection result. The method further includes: The defect report is obtained by structurally organizing the defect detection result based on the image index information. The defect report is displayed.
9. The defect detection method according to any one of claims 1 to 8, characterized in that, The target object includes a power distribution pole tower, and the original image is obtained by performing multi-view collection on the target object based on a drone, including: The original image is obtained by collecting a front view, a left view, a right view, a top view of the power distribution pole tower, and an overall view of the power distribution pole tower based on the drone.
10. A defect detection apparatus characterized by comprising: The apparatus includes: A collection module configured to obtain an original image by performing multi-view collection on a target object based on a drone; A first detection module configured to obtain a preliminary detection result by performing defect detection on the target object based on the original image; A reconstruction module configured to trigger meta-cognition processing mode in a case where the preliminary detection result indicates that a confidence that the target object has a defect is less than a preset threshold, and to obtain a three-dimensional model of the target object and collection parameter information of the original image by performing three-dimensional reconstruction based on the original image; An obtaining module configured to obtain a sampling view sequence of the target object based on the three-dimensional model, the original image, and the collection parameter information; A rendering module configured to obtain an image rendering result by rendering a three-dimensional image corresponding to the original image in the three-dimensional model based on the sampling view sequence; A second detection module configured to obtain a defect detection result by performing defect detection on the target object based on the image rendering result. 11.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein the computer program comprises instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 10. The processor implements the steps of the defect detection method in any one of claims 1-9 when executing the computer program.
12. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, which is executed by a processor, implements the steps of the defect detection method according to any one of claims 1-9.
Citation Information
Patent Citations
Aircraft skin defect identification and positioning method based on unmanned aerial vehicle and multi-view geometry
CN114937134A
Three-dimensional reconstruction method of target object, electronic equipment and storage medium
CN118823265A