An unmanned aerial vehicle inspection system and method based on air-ground cooperation and AI visual detection
Patent Information
- Application Number
- CN202610716837.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-11
AI Technical Summary
针对缺陷检测,部分方案引入基于深度学习的视觉检测或分割网络对离线数据进行分析,但多以单帧或短片段处理为主,缺乏对连续视频序列中的背景稳定结构、周期性扰动以及跨帧一致性的统一建模,导致在存在光照闪烁、旋翼阴影、纹理干扰、视角变化或目标紧邻场景时,容易出现误检、漏检、实例黏连及告警抖动问题;由于判读发生在采集之后,无法在巡检过程中根据识别结果即时调整飞行姿态、云台角度、拍摄距离或补拍策略,进而造成疑似缺陷证据不足、返航后需二次复飞取证、巡检周期延长和返工率上升
本发明通过地空协同任务下发、机载在线处理与AI视觉检测的联合设计,提升了无人机巡检的实时识别与取证联动能力。与现有按预设航线采集后离线判读的方式相比,本发明利用块循环低秩卷积分解与复数域相位共享分解对视频序列中的亮度结构与周期扰动进行分离,获得稀疏异常分量并生成环形候选带,使后续缺陷分析聚焦于有效区域;改进PSENet网络通过隐式距离场核显式化与向量流同伦归属实现精细分割与实例归属,并结合跨帧重力势能映射保持实例一致,从而在复杂光照、抖动与紧邻目标场景下提高缺陷检出率与稳定性,降低漏检与返工率。
Smart Images

Figure CN122551223A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AI visual inspection technology, and in particular to a drone inspection system and method based on ground-air collaboration and AI visual inspection. Background Technology
[0002] Drone inspections are widely used for routine inspections and defect detection of power lines, substations, photovoltaic fields, utility tunnels, and bridges. Current technologies typically involve pre-planning flight paths from the ground and transmitting them to the drone. The drone then completes its flight and takes pictures according to preset waypoints, and the collected images or videos are transmitted back to the ground for offline interpretation or manual review. For defect detection, some solutions introduce deep learning-based visual detection or segmentation networks to analyze offline data, but these primarily focus on single-frame or short-segment processing, lacking unified modeling of stable background structures, periodic disturbances, and cross-frame consistency in continuous video sequences. This leads to issues such as false positives, missed positives, instance overlap, and alarm jitter when there are lighting flicker, rotor shadows, texture interference, changes in perspective, or targets closely adjacent to the scene. Furthermore, since interpretation occurs after data acquisition, it's impossible to adjust flight attitude, gimbal angle, shooting distance, or reshooting strategies in real-time based on the recognition results during the inspection process. This results in insufficient evidence of suspected defects, requiring a second flight for evidence collection after returning to base, extended inspection cycles, and increased rework rates.
[0003] Existing UAV inspection systems, under conditions of limited air-to-ground communication bandwidth, fluctuating link latency, or severe packet loss, typically employ fixed backhaul methods and fixed processing locations. They lack hierarchical encapsulation and scheduling mechanisms oriented towards link status, leading to alarm information, clipped evidence, and original video competing for bandwidth in the same transmission channel. Critical alarms may not arrive first or, upon arrival, lack verifiable evidence fragments. During short-term link interruptions, some solutions can only perform simple buffering or passive retransmission, lacking backhaul queue management and retransmission order control tied to task status, easily resulting in fragment loss, disordered order, or duplicate backhauls. Existing solutions lack a unified structured result representation and mapping relationship between the airborne and ground ends, making it difficult for the ground end to aggregate information on the same defect across different frames, perspectives, or re-flight batches, hindering the formation of stable defect instance links and traceable evidence chains.
[0004] Therefore, how to provide a drone inspection system and method based on ground-air collaboration and AI visual inspection is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a UAV inspection system and method based on ground-air collaboration and AI visual detection. This invention comprehensively utilizes ground-air task collaboration, link adaptive hierarchical backhaul, low-rank matrix binary decomposition, and improved PSENet network technology. It details the entire process from generating and issuing inspection tasks from the ground end, to the UAV acquiring and preprocessing video, obtaining sparse anomaly components through block cyclic low-rank convolution decomposition and complex domain phase-sharing decomposition to generate ring candidate bands, achieving defect segmentation and instance consistency using implicit range field kernel explicitization, vector flow homotopy attribution, and cross-frame gravitational potential energy mapping, mapping the defect results to a three-dimensional track coordinate system to generate structured inspection results, and then hierarchically encapsulating and backhauling the results based on link status to generate and archive inspection reports. Compared with existing technologies, this invention can achieve online defect identification and evidence collection linkage during the inspection process, improve alarm timeliness and backhaul efficiency under weak network conditions, and enhance robustness and engineering feasibility under complex lighting and periodic disturbance scenarios.
[0006] A drone inspection method based on ground-air collaboration and AI visual detection according to an embodiment of the present invention includes: Inspection tasks are generated on the ground and sent to the drone. The drone performs flight and shooting according to the inspection tasks and collects video data. The video data is preprocessed to form data to be analyzed. Perform low-rank matrix binary decomposition, rearrange the data to be analyzed into a block cyclic matrix based on the block cyclic low-rank convolution solution, complete the calculation of low-rank components in the block cyclic domain, and use complex domain phase-shared decomposition to represent the data to be analyzed as complex domain factors, and map the amplitude components and phase components into brightness structure and periodic perturbation respectively to obtain sparse anomaly components. A pixel displacement vector field is established based on sparse anomaly components. The divergence distribution of the pixel displacement vector field is calculated. Anomaly candidate centers are marked at local maxima of divergence. A ring-shaped candidate band is generated by continuously tracking along the divergence gradient. An improved PSENet network is constructed to perform defect segmentation and instance attribution processing on the annular candidate band. Implicit distance field kernels are introduced to make the continuous distance field prediction explicit. Two-dimensional vector flow prediction is performed based on vector flow homotopy attribution. Pixels are merged into the corresponding kernels along the vector flow trajectory. Potential energy matching is performed using cross-frame gravitational potential energy mapping. Defect instance characterization data is output. The defect instance characterization data is mapped to the current three-dimensional trajectory coordinate system of the UAV, a dynamic surface is constructed, and continuous deduction is performed in the tangent direction of the dynamic surface to generate the structured inspection results. Based on link status parameters and backhaul priority parameters, the inspection structured structure is hierarchically encapsulated and backhaul scheduled to generate an inspection report and complete archiving.
[0007] Optionally, the inspection task includes a sequence of waypoints arranged in chronological order, the flight altitude, flight speed, heading angle range, gimbal pitch angle range, zoom range, and shooting time markers corresponding to each waypoint.
[0008] Optionally, the preprocessing of video data to form data to be analyzed includes: On the ground, inspection tasks are generated based on the geographical range of the target to be inspected, the boundary of the no-fly zone, the coordinates of the take-off and landing points, and the operation time window. The inspection tasks are then encoded into task data packets. The mission data packet is sent to the UAV terminal through the air-to-ground communication link. The UAV terminal performs integrity verification and parsing of the mission data packet, establishes the mapping relationship between mission identifier and waypoint sequence, writes the waypoint sequence into the flight control queue, and writes the shooting time identifier and gimbal control parameters into the payload control queue. The UAV arrives at each waypoint in sequence according to the flight control queue and performs shooting according to the payload control queue. It collects video data and synchronously records the timestamp, positioning coordinates, attitude angle and gimbal parameters corresponding to each frame of video data. The video data is preprocessed to form data to be analyzed. The preprocessing includes frame extraction at a fixed frame rate, resolution unification of the extracted frames, geometric alignment of the extracted frames, and assembling the extracted frames into time-series data segments according to the number of consecutive frames.
[0009] Optionally, obtaining the sparse outlier components includes: The data to be analyzed is constructed into an observation matrix in the order of consecutive frames. The observation matrix is formed by stitching multiple frames of images together column by column. Each column is a pixel vector obtained by expanding a frame of images in row priority order. The observation matrix is divided into sub-block arrays according to the sub-block size. Based on the block-circular low-rank convolution solution, the sub-block array is rearranged in a block-circular manner to form a block-circular matrix. The block-circular matrix satisfies that the sub-blocks in the next row are cyclically shifted relative to the sub-blocks in the previous row by a fixed number of sub-blocks. The low-rank factor is iteratively updated in the block-circular domain to obtain the low-rank factor matrix and the low-rank coefficient matrix. The low-rank background component is obtained by multiplying the transpose of the low-rank factor matrix and the low-rank coefficient matrix. The residual matrix is obtained by subtracting the observation matrix from the low-rank background component. The residual matrix is represented as a complex domain residual matrix by using complex domain phase-sharing decomposition. The complex domain residual matrix is decomposed into a complex domain low-rank periodic component and a complex domain sparse anomaly component. The complex domain low-rank periodic component is obtained by multiplying the complex domain basis matrix and the transpose of the complex domain coefficient matrix. Each element of the complex domain basis matrix and the complex domain coefficient matrix consists of an amplitude component and a phase component. The amplitude component represents the brightness structure in the residual, and the phase component represents the periodic perturbation in the residual. The sparse outlier component in the complex field is obtained by the difference between the observation matrix and the low-rank background component in the complex field. The absolute value operation of the amplitude of the sparse outlier component in the complex field is performed to obtain the sparse outlier component. The amplitude of each element in the sparse outlier component is equal to the square root of the sum of the squares of the real part and the squares of the imaginary part of the element.
[0010] Optionally, generating the annular candidate band includes: Using the sparse outlier component corresponding to each frame of two-dimensional pixel array as input, spatial gradient calculation is performed on the two-dimensional pixel array to obtain the horizontal gradient field and the vertical gradient field. The horizontal gradient field and the vertical gradient field at each pixel position form the pixel displacement vector field. The vector of the pixel displacement vector field at each pixel position is composed of the horizontal component and the vertical component. The divergence distribution is calculated for the pixel displacement vector field. The divergence value at any pixel position is equal to the sum of the partial derivatives of the horizontal component of the pixel position with respect to the horizontal coordinate and the partial derivatives of the vertical component with respect to the vertical coordinate, forming a divergence map. Local maxima in the divergence map are identified and used as candidate anomaly centers. A closed trajectory is generated by stepping outward from the candidate anomaly centers along the gradient direction of the divergence map. The internal region of the closed trajectory is merged with the neighboring region of the closed trajectory to form an annular candidate band.
[0011] Optionally, the output defect instance characterization data includes: An improved PSENet network is constructed and the pixel features corresponding to the annular candidate band are input. The improved PSENet network sets an implicit range field output branch at the segmentation output position of the original PSENet network to make the implicit range field kernel explicit. A two-dimensional vector flow is set at the pixel-level prediction output position of the original PSENet network to perform vector flow homotopy assignment. After the instance generation output, cross-frame gravitational potential energy mapping is performed at the cross-frame identifier mapping position. The continuous range field obtained from the implicit range field output branch is kernel-explicitly processed based on the implicit range field kernel explicitization, generating multiple sets of range level sets and generating a kernel set from the pixel set corresponding to each range level set. The range level set is the set of pixels that satisfy the distance value equals the distance constant. Based on the homotopy assignment of vector flow, the two-dimensional vector flow obtained from the output branch of the two-dimensional vector flow is processed for instance assignment. Starting from each pixel, the two-dimensional vector flow is iteratively updated along the two-dimensional vector flow to obtain the pixel trajectory. The pixel coordinates are updated to the sum of the two-dimensional vectors at the previous updated pixel coordinates and the current pixel coordinates until the pixel trajectory enters any kernel region in the kernel set. The kernel identifier corresponding to the entered kernel region is assigned as the defect instance identifier of the pixel to generate the defect instance mask of the current frame. Cross-frame gravity potential energy mapping is used to perform cross-frame consistency processing on the defect instance identifier of the current frame. The potential energy contribution is calculated at the centroid position of each defect instance mask in the previous frame and the potential energy value is calculated at the centroid position of the defect instance in the current frame. The defect instance identifier of the previous frame corresponding to the minimum potential energy value is assigned as the defect instance identifier of the current frame, and the defect instance characterization data is output. The improved PSENet network is trained by constructing training samples and updating network parameters. The training samples include circular candidate band images and corresponding annotations. The network parameters are updated using the total loss, which is equal to the sum of the segmentation loss corresponding to the pixel annotation of the defect region, the distance field loss corresponding to the continuous distance field annotation, the vector flow loss corresponding to the two-dimensional vector flow annotation, and the consistency loss corresponding to the defect instance identifier annotation of the adjacent frame.
[0012] Optionally, generating the structured inspection results includes: Each defect instance mask in the defect instance characterization data is converted into a set of pixel coordinates. The centroid coordinates of the pixel coordinate set are calculated and the boundary coordinate set of the defect instance is determined. The geometric quantization parameters of the defect instance are calculated based on the boundary coordinate set. Based on timestamps, positioning data, pose data and gimbal parameters, establish a mapping relationship from pixel coordinates to a three-dimensional track coordinate system, map the centroid coordinates and boundary coordinates set to the defect center coordinates and defect boundary coordinates set, and associate defect instance identifiers and evidence indexes to the defect center coordinates and defect boundary coordinates set. A dynamic surface is constructed based on the set of defect center coordinates and defect boundary coordinates. The dynamic surface is defined by the current position of the UAV, the range of allowable flight altitude, the range of allowable pitch angle, and the range of safe distance. The reshoot state sequence is continuously generated in the tangent direction of the dynamic surface according to the step size. The reshoot state sequence and defect instance characterization data are jointly encapsulated to generate the inspection structured result.
[0013] Optionally, the step of performing hierarchical encapsulation and backhaul scheduling, generating inspection reports, and completing archiving includes: Collect link status parameters of the air-to-ground communication link and generate link status identifiers. The link status parameters include uplink bandwidth, downlink bandwidth, round-trip time, and packet loss rate. The inspection structured results are divided into alarm summary data, evidence trimming data and raw data according to data type, and data packages are generated for each. A backhaul queue is established based on the backhaul priority parameter and scheduled in conjunction with the link status identifier. The backhaul queue is arranged in the order of alarm summary data, evidence clipping data, and original data. Data packets are backhauled sequentially within the allowed backhaul range corresponding to the link status identifier, and data packets exceeding the allowed backhaul range are buffered. When the allowed backhaul range corresponding to the link status identifier expands, the buffered data packets are supplemented according to the backhaul queue order. The ground end verifies and unpacks the received data packets and generates an inspection report for archiving.
[0014] According to an embodiment of the present invention, a drone inspection system based on ground-air collaboration and AI visual inspection includes the following modules: The task assignment module is used to generate inspection tasks on the ground, drive the UAV to complete flight shooting and video acquisition preprocessing, and form data to be analyzed; The low-rank decomposition module is used to perform low-rank matrix binary decomposition on the data to be analyzed and output sparse outlier components. The candidate band generation module is used to establish a pixel displacement vector field based on sparse anomaly components and calculate the divergence distribution, locate the anomaly candidate center and generate a ring-shaped candidate band. An improved segmentation module is used to construct an improved PSENet network and perform continuous range field prediction, two-dimensional vector flow prediction and cross-frame potential matching on the annular candidate band, outputting defect instance characterization data. The structured generation module is used to map defect instance characterization data to a three-dimensional track coordinate system, construct dynamic surface deduction, and generate structured inspection results.
[0015] The hierarchical feedback module is used to hierarchically encapsulate and schedule the feedback of the structured inspection results, generate inspection reports, and complete archiving.
[0016] The beneficial effects of this invention are: This invention enhances the real-time identification and evidence collection capabilities of UAV inspections through a joint design of ground-air collaborative task deployment, airborne online processing, and AI visual inspection. Compared with existing methods that collect data along preset routes and then interpret it offline, this invention utilizes block cyclic low-rank convolution decomposition and complex domain phase-sharing decomposition to separate the brightness structure and periodic perturbations in video sequences, obtaining sparse anomalous components and generating ring candidate bands, allowing subsequent defect analysis to focus on effective regions. The improved PSENet network achieves fine segmentation and instance attribution through implicit distance field kernel explicitization and vector flow homotopy attribution, and combines cross-frame gravitational potential energy mapping to maintain instance consistency, thereby improving defect detection rate and stability under complex lighting, jitter, and adjacent target scenarios, and reducing missed detection and rework rates.
[0017] This invention achieves efficient backhaul and structured delivery under weak network conditions through link status awareness and backhaul priority control. Compared to existing schemes with fixed backhaul granularity and fixed processing locations, this invention maps defect instance characterization data to a three-dimensional track coordinate system and constructs dynamic surfaces to generate structured inspection results, enabling defect locations, geometric quantification parameters, and evidence indexes to form a traceable data chain. The structured inspection results are hierarchically encapsulated and backhaul scheduled, prioritizing the delivery of alarm summaries and key evidence and supporting cached retransmission, improving alarm timeliness, data transmission efficiency, and report generation automation, facilitating closed-loop management with operation and maintenance processes. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a drone inspection method based on ground-air collaboration and AI visual detection proposed in this invention; Figure 2 This is a block diagram of the improved PSENet network for a drone inspection method based on ground-air cooperation and AI visual detection proposed in this invention. Figure 3 This is a functional diagram of a drone inspection system based on ground-air collaboration and AI visual inspection proposed in this invention. Detailed Implementation
[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0020] refer to Figure 1 and Figure 2 A drone inspection method based on ground-air collaboration and AI visual inspection includes: Inspection tasks are generated on the ground and sent to the drone. The drone performs flight and shooting according to the inspection tasks and collects video data. The video data is preprocessed to form data to be analyzed. Perform low-rank matrix binary decomposition, rearrange the data to be analyzed into a block cyclic matrix based on the block cyclic low-rank convolution solution, complete the calculation of low-rank components in the block cyclic domain, and use complex domain phase-shared decomposition to represent the data to be analyzed as complex domain factors, and map the amplitude components and phase components into brightness structure and periodic perturbation respectively to obtain sparse anomaly components. A pixel displacement vector field is established based on sparse anomaly components. The divergence distribution of the pixel displacement vector field is calculated. Anomaly candidate centers are marked at local maxima of divergence. A ring-shaped candidate band is generated by continuously tracking along the divergence gradient. An improved PSENet network is constructed to perform defect segmentation and instance attribution processing on the annular candidate band. Implicit distance field kernels are introduced to make the continuous distance field prediction explicit. Two-dimensional vector flow prediction is performed based on vector flow homotopy attribution. Pixels are merged into the corresponding kernels along the vector flow trajectory. Potential energy matching is performed using cross-frame gravitational potential energy mapping. Defect instance characterization data is output. The defect instance characterization data is mapped to the current three-dimensional trajectory coordinate system of the UAV, a dynamic surface is constructed, and continuous deduction is performed in the tangent direction of the dynamic surface to generate the structured inspection results. Based on link status parameters and backhaul priority parameters, the inspection structured structure is hierarchically encapsulated and backhaul scheduled to generate an inspection report and complete archiving.
[0021] In this embodiment, the inspection task includes a waypoint sequence arranged in chronological order, the flight altitude, flight speed, heading angle range, gimbal pitch angle range, zoom range, and shooting time marker for each waypoint.
[0022] In this embodiment, the preprocessing of video data to form data to be analyzed includes: On the ground, inspection tasks are generated based on the geographical range of the target to be inspected, the boundary of the no-fly zone, the coordinates of the take-off and landing points, and the operation time window. The inspection tasks are then encoded into task data packets. The mission data packet is sent to the UAV via the air-to-ground communication link. The UAV performs integrity verification and parsing of the mission data packet, establishes a mapping relationship between the mission identifier and the waypoint sequence, writes the waypoint sequence into the flight control queue, and writes the shooting time identifier and gimbal control parameters into the payload control queue. Specifically, the UAV performs integrity verification and parsing of the mission data packet as follows: The payload length field is read and the actual number of received bytes is checked to ensure it matches the payload length. A 16-bit cyclic redundancy check value is calculated sequentially for all bytes except the check field. The calculation result is compared with the 16-bit check value carried at the end of the data packet. If they do not match, the packet is discarded and a retransmission request is sent. If they match, the packet is parsed. During parsing, the task identifier, number of waypoints, waypoint sequence, shooting time identifier, and gimbal control parameters are extracted in the order of the fields. The altitude, speed, angle, and zoom ratio are checked for ranges of 5 meters to 200 meters, 0.1 meters per second to 20 meters per second, -180 degrees to +180 degrees with the lower limit less than the upper limit, and zoom ratio from 1x to the maximum payload ratio. After all checks are passed, a mapping between the task identifier and the waypoint sequence is established. The waypoint sequence is written to the flight control queue, and the shooting time identifier and gimbal control parameters are written to the payload control queue. The UAV arrives at each waypoint in sequence according to the flight control queue and performs shooting according to the payload control queue. It collects video data and synchronously records the timestamp, positioning coordinates, attitude angle and gimbal parameters corresponding to each frame of video data. The video data is preprocessed to form data to be analyzed. The preprocessing includes frame extraction at a fixed frame rate, resolution unification of the extracted frames, geometric alignment of the extracted frames, and assembling the extracted frames into time-series data segments according to the number of consecutive frames.
[0023] In this embodiment, obtaining the sparse outlier components includes: The data to be analyzed is constructed into an observation matrix in consecutive frame order. The observation matrix is formed by concatenating multiple frames of images column by column. Each column is a pixel vector obtained by expanding one frame of image in row-major order. The observation matrix is then divided into sub-block arrays according to the sub-block size, where: The data to be analyzed is constructed into an observation matrix in the order of consecutive frames. Specifically, from a time series data segment of the data to be analyzed, consecutive frames are selected in ascending order of timestamp. Each frame is unified to the same resolution and converted into a grayscale image with a grayscale value range of 0 to 255. Each frame is expanded into a pixel vector in row-major order. The expansion order is to first take all pixels from left to right in the first row, then take all pixels from left to right in the second row, and so on until the last row. The pixel vector of the first frame is used as the first column, the pixel vector of the second frame is used as the second column, and the pixel vectors of the remaining frames are concatenated column by column to obtain the observation matrix. The observation matrix is divided into a sub-block array according to the sub-block size. Specifically, the pixel vector corresponding to each column in the observation matrix is restored to a two-dimensional frame image according to the original row and column positions. Each frame image is divided into non-overlapping blocks of 16×16 pixels. The block division order is to divide the blocks horizontally from left to right, starting from the top left corner, and then vertically from top to bottom. Each sub-block is expanded into a sub-block vector in row priority order. The sub-block vectors of the same spatial position in consecutive frames are concatenated column by column to form a sub-block matrix. All sub-block matrices are arranged according to their row and column positions in the frame image to form a sub-block array. When the height or width of the frame image is not divisible by 16, the bottom and right boundaries are filled with pixel values of 0 to make the height and width integer multiples of 16 before block division is performed. Based on the block-cyclic low-rank convolution solution, the sub-block array is cyclically rearranged to form a block-cyclic matrix. The block-cyclic matrix satisfies the condition that each row of sub-blocks is cyclically shifted relative to the previous row by a fixed number of sub-blocks. The low-rank factors are iteratively updated in the block-cyclic domain to obtain the low-rank factor matrix and the low-rank coefficient matrix. The low-rank background component is obtained by multiplying the transposes of the low-rank factor matrix and the low-rank coefficient matrix. The residual matrix is obtained by subtracting the observation matrix from the low-rank background component. Where: The sub-block array is cyclically rearranged to form a cyclic matrix. Specifically, the sub-block array is fixed into a two-dimensional grid with the number of rows equal to the number of vertical sub-blocks and the number of columns equal to the number of horizontal sub-blocks. Each grid cell stores one sub-block matrix. Taking the first row as the base row, the second row is cyclically shifted one sub-block position to the right relative to the first row. The third row is then cyclically shifted one sub-block position to the right relative to the second row, and so on until the last row. Cyclicly shifting one sub-block position to the right means moving the rightmost sub-block of the current row to the leftmost row, while shifting the remaining sub-blocks one position to the right without changing the data inside the sub-blocks. After completing the cyclic shift of each row, the sub-blocks are concatenated in row priority order to form a cyclic matrix. The concatenation order is as follows: first, all sub-block matrices of the first row from left to right are concatenated sequentially, then all sub-block matrices of the second row from left to right are concatenated sequentially, until the last row. The low-rank factor matrix and low-rank coefficient matrix are obtained by iteratively updating the low-rank factor in the block cyclic domain. Specifically, the block cyclic matrix is transformed into a block cyclic domain representation using a two-dimensional discrete Fourier transform, and the block sequence at each sub-block position is transformed to obtain a frequency domain block matrix. The low-rank factor matrix and low-rank coefficient matrix are initialized on the frequency domain block matrix. The initialization method is that the elements of the low-rank factor matrix are uniformly random numbers between 0 and 1, and the elements of the low-rank coefficient matrix are 0. Iterative updates are performed using alternating least squares. First, the low-rank factor matrix is fixed and the low-rank coefficient matrix is updated using alternating least squares. Then, the updated low-rank coefficient matrix is fixed and the low-rank factor matrix is updated using alternating least squares. The two-step update is considered as one iteration and is repeated 20 times. After 20 iterations, a two-dimensional discrete Fourier inverse transform is performed to obtain the spatial domain low-rank factor matrix and low-rank coefficient matrix. The residual matrix is represented as a complex-domain residual matrix using complex-domain phase-shared decomposition. This complex-domain residual matrix is then decomposed into a low-rank complex-domain periodic component and a sparse anomaly component. The low-rank complex-domain periodic component is obtained by multiplying the transpose of the complex-domain basis matrix and the complex-domain coefficient matrix. Each element of the complex-domain basis matrix and the complex-domain coefficient matrix consists of an amplitude component and a phase component. The amplitude component represents the brightness structure in the residual, and the phase component represents the periodic perturbation in the residual. Wherein: Representing the residual matrix as a complex domain residual matrix involves: preserving the time order of the columns of the residual matrix, restoring each column to a two-dimensional residual frame and forming a time series at the same pixel position, performing a discrete Fourier transform on the time series at each pixel position, taking the frequency component with the largest amplitude in the transform result as the periodic component, and calculating the amplitude and phase of the frequency component. The amplitude is equal to the square root of the sum of the squares of the real part and the squares of the imaginary part, and the phase is equal to the imaginary part of the arctangent function divided by the real part. The amplitude at each pixel position is used as the real part of the complex number, and the phase is used as the imaginary part of the complex number to construct complex pixel values. The complex pixel values of each frame are expanded into complex vectors in row-major order and concatenated column by column to obtain the complex domain residual matrix. The complex domain residual matrix is decomposed into a low-rank periodic component and a sparse outlier component in the complex domain. Specifically, the complex domain basis matrix and the complex domain coefficient matrix are initialized in the complex domain. The elements of the basis matrix and the coefficient matrix are composed of magnitude components and phase components. During initialization, the magnitude components are uniformly random numbers between 0 and 1, and the phase components are uniformly random numbers between negative pi and positive pi. Alternating least squares iterative updates are used. After the update is completed, the low-rank periodic component in the complex domain is obtained by multiplying the transpose of the complex domain basis matrix and the complex domain coefficient matrix. The sparse outlier component in the complex domain is obtained by the difference between the complex domain residual matrix and the low-rank periodic component in the complex domain. The sparse outlier component in the complex field is obtained by the difference between the observation matrix and the low-rank background component in the complex field. The absolute value operation of the amplitude of the sparse outlier component in the complex field is performed to obtain the sparse outlier component. The amplitude of each element in the sparse outlier component is equal to the square root of the sum of the squares of the real part and the squares of the imaginary part of the element.
[0024] In this embodiment, generating the annular candidate band includes: Using the 2D pixel array corresponding to each frame of sparse outlier components as input, spatial gradient calculation is performed on the 2D pixel array to obtain the horizontal gradient field and the vertical gradient field. The horizontal and vertical gradient fields at each pixel position form a pixel displacement vector field, and the vector at each pixel position is composed of horizontal and vertical components. Specifically, the spatial gradient calculation of the 2D pixel array to obtain the horizontal and vertical gradient fields is as follows: The gradient of the two-dimensional pixel array is calculated in the spatial domain using a 3x3 Sobel operator, and the pixel values of the two-dimensional pixel array are normalized to 0 to 1. The horizontal gradient field is obtained by convolving the 3x3 neighborhood of each pixel with a horizontal convolution kernel. The horizontal convolution kernel has values of -1, 0, and +1 in the first row, -2, 0, and +2 in the second row, and -1, 0, and +1 in the third row. The vertical gradient field is obtained by convolving the 3x3 neighborhood of each pixel with a vertical convolution kernel. The vertical convolution kernel has values of -1, -2, and -1 in the first row, 0, 0, and +1 in the second row, and +1, +2, and +1 in the third row. The boundary of the two-dimensional pixel array is filled with mirror image. The convolution results are used as the horizontal gradient value and the vertical gradient value of the pixel position, respectively. The horizontal gradient value and the vertical gradient value constitute the horizontal component and the vertical component of the pixel displacement vector of the pixel position. The divergence distribution of the pixel displacement vector field is calculated. The divergence value at any pixel location is equal to the sum of the partial derivatives of the horizontal component of the pixel location with respect to the horizontal coordinate and the partial derivatives of the vertical component with respect to the vertical coordinate, forming a divergence map. Specifically, the calculation of the divergence distribution of the pixel displacement vector field is as follows: For each non-boundary pixel, first calculate the partial derivative of the horizontal component in the horizontal coordinate. The partial derivative is equal to the horizontal component of the right adjacent pixel minus the horizontal component of the left adjacent pixel and then divided by 2. Then calculate the partial derivative of the vertical component in the vertical coordinate. The partial derivative is equal to the vertical component of the lower adjacent pixel minus the vertical component of the upper adjacent pixel and then divided by 2. Add the two partial derivatives to get the divergence value of the pixel position. For boundary pixels, the partial derivatives are calculated using forward and backward differencing. The leftmost column horizontal partial derivative is equal to the horizontal component of the adjacent pixel to the right minus the horizontal component of the current pixel. The rightmost column horizontal partial derivative is equal to the horizontal component of the current pixel minus the horizontal component of the adjacent pixel to the left. The top row vertical partial derivative is equal to the vertical component of the adjacent pixel below minus the vertical component of the current pixel. The bottom row vertical partial derivative is equal to the vertical component of the current pixel minus the vertical component of the adjacent pixel above. The divergence values of all pixel positions are combined to form a divergence map. Local maxima in the divergence map are identified and used as candidate anomaly centers. A closed trajectory is generated by stepping outward from the candidate anomaly centers along the gradient direction of the divergence map. The internal region of the closed trajectory is merged with the neighboring region of the closed trajectory to form an annular candidate band.
[0025] In this embodiment, the output defect instance characterization data includes: An improved PSENet network is constructed and the pixel features corresponding to the annular candidate band are input. The improved PSENet network sets an implicit range field output branch at the segmentation output position of the original PSENet network to make the implicit range field kernel explicit. A two-dimensional vector flow is set at the pixel-level prediction output position of the original PSENet network to perform vector flow homotopy assignment. After the instance generation output, cross-frame gravitational potential energy mapping is performed at the cross-frame identifier mapping position. The continuous range field obtained from the implicit range field output branch is explicitly kernelized based on implicit range field kernelization. This generates multiple sets of range level sets, and the pixel sets corresponding to each range level set are used to generate a kernel set. A range level set is a set of pixels whose distance value equals a distance constant. Specifically, the kernelization of the continuous range field obtained from the implicit range field output branch is performed as follows: The continuous distance field obtained from the implicit distance field output branch is first normalized by range normalization. The minimum distance value in the distance field is linearly mapped to 0 and the maximum distance value is linearly mapped to 1. Six distance constants are set, namely 0.1, 0.2, 0.3, 0.4, 0.5 and 0.6. For each distance constant, a corresponding distance level set pixel set is generated. The condition for the pixel set is that the absolute value of the difference between the normalized distance value of the pixel and the corresponding distance constant is not greater than 0.02. Eight-neighbor connectivity aggregation is performed on the pixel set obtained for the same distance constant to form several connected regions. All connected regions corresponding to the distance constant are arranged in ascending order of distance constant and merged into a kernel set. Based on the homotopy assignment of vector streams, instance assignment processing is performed on the two-dimensional vector stream obtained from the output branch of the two-dimensional vector stream. Starting from each pixel, the process iteratively updates the pixel trajectory along the two-dimensional vector stream, updating the pixel coordinates to the sum of the two-dimensional vectors at the previously updated pixel coordinates and the current pixel coordinates. This continues until the pixel trajectory enters any kernel region in the kernel set. The kernel identifier corresponding to the entered kernel region is assigned as the defect instance identifier of the pixel, generating a defect instance mask for the current frame. Specifically, the instance assignment processing for the two-dimensional vector stream obtained from the output branch of the two-dimensional vector stream is as follows: The two-dimensional vector stream obtained from the output branch of the two-dimensional vector stream is first subjected to amplitude truncation processing, with a truncation upper limit of 1.0 pixels. The horizontal and vertical components of the two-dimensional vector at any pixel position are restricted to -1.0 to +1.0 respectively. Trajectory iteration and assignment are performed for each pixel. The iteration start point is the initial coordinate of the pixel, and the iteration step size is 1.0. In each iteration, the current coordinate is updated to the current coordinate plus the two-dimensional vector at the coordinate. The updated horizontal and vertical coordinates are taken as the nearest integer pixel coordinates. The maximum number of iterations for each pixel is set to 30. When the updated coordinate falls into the pixel set of any kernel layer in the kernel set, the iteration stops. The kernel identifier corresponding to the kernel is assigned as the defect instance identifier of the pixel. When it has been iterated for 30 times and still has not entered any kernel pixel set, the defect instance identifier is assigned as 0 to indicate that it has not been assigned. The defect instance identifiers of all pixels in the frame are arranged according to the pixel position to form the defect instance mask of the current frame. Cross-frame consistency processing is performed on the defect instance identifiers of the current frame using cross-frame gravity potential energy mapping. The potential energy contribution is calculated at the centroid position of each defect instance mask in the previous frame, and the potential energy value is calculated at the centroid position of the defect instance in the current frame. The defect instance identifier of the previous frame corresponding to the minimum potential energy value is assigned as the defect instance identifier of the current frame, and the defect instance characterization data is output. Specifically, the cross-frame consistency processing for the defect instance identifiers of the current frame involves: Centroid coordinates are calculated for each defect instance mask in the previous and current frames. The centroid coordinates are equal to the arithmetic mean of the horizontal and vertical coordinates of all pixels in the instance mask. For each defect instance in the current frame, the potential energy value between it and each defect instance in the previous frame is calculated. The potential energy value is equal to the sum of the potential energy contributions of the defect instances in the previous frame to the defect instances in the current frame. The potential energy contribution is calculated as 1 divided by the distance, which is equal to the Euclidean distance between the centroids of the current frame and the centroids of the previous frame plus 0.001. After calculating all potential energy values, the defect instance identifier of the previous frame corresponding to the smallest potential energy value is selected as the new identifier of the defect instance in the current frame. When the minimum potential energy value is greater than 1000, the original identifier of the current frame is kept unchanged. The defect instance identifier after the identifier replacement is written back to the defect instance mask in the current frame. The defect instance identifier, defect instance mask, centroid coordinates, boundary coordinate set and evidence index are merged to generate defect instance representation data. The improved PSENet network is trained by constructing training samples and updating network parameters. The training samples include images of annular candidate bands and their corresponding annotations. The network parameters are updated using a total loss, which is equal to the sum of the segmentation loss corresponding to the pixel annotations of the defect region, the distance field loss corresponding to the continuous distance field annotations, the vector flow loss corresponding to the two-dimensional vector flow annotations, and the consistency loss corresponding to the defect instance identifiers of adjacent frames. The segmentation loss corresponding to the pixel annotations in the defect region and the range field loss corresponding to the continuous range field annotations are as follows: The segmentation loss is the sum of pixel-level binary cross-entropy loss and dice loss. For the binary cross-entropy loss, the negative logarithm of the predicted probability is taken when the pixel is labeled as 1 and the negative logarithm of the predicted probability is taken when the pixel is labeled as 0. The arithmetic mean is taken for all pixels. The dice loss first multiplies the predicted probability map and the pixel label map of the defect area pixel by pixel and sums them to obtain the intersection sum. Then, the predicted probability map and the label map are summed pixel by pixel to obtain the union sum. The dice coefficient is equal to 2 multiplied by the intersection sum plus 1 and divided by the union sum plus 1. The dice loss is equal to 1 minus the dice coefficient. The distance field loss adopts pixel-level absolute error loss. For each pixel, the absolute value of the difference between the predicted distance value and the labeled distance value is calculated and the arithmetic mean is taken. The labeled distance value is generated by the pixel label of the defect area. The generation method is to take the minimum Euclidean distance from the defect boundary to the pixel inside the defect area and the minimum Euclidean distance from the defect boundary to the pixel outside the defect area and take the negative sign. The vector flow loss corresponding to the two-dimensional vector flow annotation and the consistency loss corresponding to the defect instance identifier annotation in adjacent frames are as follows: The vector flow loss employs a pixel-wise smoothing absolute error loss method. The target vector for each pixel is given by the two-dimensional vector flow annotation. The target vector points to the nearest pixel in the kernel to which the pixel belongs. The horizontal component of the target vector is equal to the x-coordinate of the nearest point in the kernel minus the x-coordinate of the current pixel, and the vertical component is equal to the y-coordinate of the nearest point in the kernel minus the y-coordinate of the current pixel. The magnitude of the target vector is truncated to an upper limit of 1.0 pixels. The smoothing absolute error loss is calculated separately for each component. When the absolute value of the difference between the predicted component and the labeled component is not greater than 1.0, 0.5 is multiplied by the square of the difference. When the absolute value is greater than 1.0, the absolute value of the difference is subtracted from 0.5. The horizontal and vertical components are summed and then the arithmetic mean is taken over all pixels. The consistency loss adopts the cross-frame mask consistency loss. First, according to the cross-frame gravitational potential energy mapping, each defect instance identifier in the current frame is replaced with the corresponding defect instance identifier in the previous frame. Then, the overlap of the same identifier in the two frames is calculated. The overlap is equal to the number of intersection pixels divided by the number of union pixels. The consistency loss is equal to 1 minus the overlap. The arithmetic mean of the consistency loss of all defect instances in the current frame is taken. The number of intersection pixels refers to the number of pixels in both frames that are the current identifier. The number of union pixels refers to the number of pixels in either frame that are the current identifier.
[0026] In this embodiment, generating the structured inspection results includes: Each defect instance mask in the defect instance representation data is converted into a set of pixel coordinates. The centroid coordinates of the pixel coordinate set are calculated, and the boundary coordinate set of the defect instance is determined. The geometric quantization parameters of the defect instance are calculated based on the boundary coordinate set. Specifically, the calculation of the geometric quantization parameters of the defect instance based on the boundary coordinate set is as follows: The boundary coordinate set is arranged in clockwise order as a closed point sequence. Adjacent points are determined by eight-neighbor connectivity. The area parameter is calculated by pixel counting. The number of internal pixels is obtained by counting all pixels in the defect instance mask and converted to the actual area of a single pixel. The actual area of a single pixel is determined by the imaging resolution and the shooting distance. The calculation method is the width of the ground covered by a single pixel multiplied by the height of the ground covered by a single pixel. The width of the ground covered by a single pixel is equal to the shooting distance multiplied by the camera's horizontal field of view divided by the image width. The height of the ground covered by a single pixel is equal to the shooting distance multiplied by the camera's vertical field of view divided by the image height. The perimeter parameter is calculated using the boundary accumulation method. The Euclidean distance is calculated and accumulated for each pair of adjacent boundary points in the closed point sequence. The Euclidean distance is equal to the square root of the sum of the squares of the difference between the x-coordinates and the squares of the difference between the y-coordinates of the two points. The accumulated perimeter is then multiplied by the width of the ground covered by a single pixel to obtain the actual perimeter. Based on timestamps, positioning data, pose data, and gimbal parameters, a mapping relationship from pixel coordinates to a 3D track coordinate system is established. The centroid coordinates and boundary coordinates are mapped to the defect center coordinates and defect boundary coordinates. Defect instance identifiers and evidence indexes are associated with the defect center coordinates and defect boundary coordinates. Specifically, the mapping relationship from pixel coordinates to the 3D track coordinate system is established as follows: The system reads the positioning data corresponding to the current frame timestamp to obtain the UAV's position and altitude, reads the pose data to obtain the roll angle, pitch angle, and heading angle, and reads the gimbal parameters to obtain the gimbal pitch angle and gimbal yaw angle. The pixel coordinates are converted into the ray direction of the camera coordinate system according to the camera intrinsic parameters. The calculation method is to subtract the principal point's x-coordinate from the pixel x-coordinate and divide by the focal length to obtain the lateral component, and subtract the principal point's y-coordinate from the pixel y-coordinate and divide by the focal length to obtain the longitudinal component. The ray direction is taken as a normalized vector. The ray direction is transformed into the three-dimensional track coordinate system by rotating the gimbal yaw and pitch, then by heading, pitch, and roll. The three-dimensional point coordinates are obtained by finding the intersection of the ray and the ground plane with the UAV position as the starting point. The ground plane height is taken as the ground elevation. When the angle between the ray and the ground plane is less than 5 degrees, no mapping result is output. In the remaining cases, the intersection point is used as the defect center coordinates or defect boundary coordinates, and the defect instance identifier and evidence index are bound to the corresponding coordinate set. A dynamic surface is constructed based on the set of defect center coordinates and defect boundary coordinates. This dynamic surface is defined by the UAV's current position, allowable flight altitude range, allowable pitch angle range, and safe distance range. A re-enhancing state sequence is continuously generated along the tangent direction of the dynamic surface, with a step size. This re-enhancing state sequence and defect instance characterization data are then encapsulated to generate a structured inspection result. Specifically, the construction of the dynamic surface based on the set of defect center coordinates and defect boundary coordinates involves: The maximum Euclidean distance from the defect boundary coordinate set to the defect center coordinate set is used as the defect circumscribed radius. The safe distance range is set to the defect circumscribed radius plus 2 meters to the defect circumscribed radius plus 10 meters. Using the defect center coordinate set as the surface reference point, a parametric surface is constructed, consisting of altitude, horizontal distance, and line-of-sight azimuth variables. The altitude variable is taken from the lower to the upper limit of the allowable flight altitude range, the horizontal distance variable is taken from the lower to the upper limit of the safe distance range, and the line-of-sight azimuth variable is taken from 0 degrees to 360 degrees. For any parameter combination, the waypoint coordinates for reshooting are calculated as the defect center coordinates plus the displacement vector obtained by decomposing the horizontal distance in the horizontal plane according to the line-of-sight azimuth, plus the vertical displacement of the altitude variable. The gimbal attitude parameters are calculated so that the camera optical axis points to the defect center coordinates. The waypoint coordinates and gimbal attitude parameter set corresponding to all parameter combinations that satisfy the altitude range, pitch angle range, and the closest distance between the waypoint and the defect boundary is not less than the lower limit of the safe distance are defined as a dynamic surface.
[0027] In this embodiment, the step of performing hierarchical encapsulation and backhaul scheduling, generating inspection reports, and completing archiving includes: Collect link status parameters of the air-to-ground communication link and generate link status identifiers. The link status parameters include uplink bandwidth, downlink bandwidth, round-trip time, and packet loss rate. The inspection structured results are divided into alarm summary data, evidence trimming data and raw data according to data type, and data packages are generated for each. A backhaul queue is established based on the backhaul priority parameter and scheduled in conjunction with the link status identifier. The backhaul queue is arranged in the order of alarm summary data, evidence clipping data, and original data. Data packets are backhauled sequentially within the allowed backhaul range corresponding to the link status identifier, and data packets exceeding the allowed backhaul range are buffered. When the allowed backhaul range corresponding to the link status identifier expands, the buffered data packets are supplemented according to the backhaul queue order. The ground end verifies and unpacks the received data packets and generates an inspection report for archiving.
[0028] refer to Figure 3 A drone inspection system based on ground-air collaboration and AI visual inspection includes the following modules: The task assignment module is used to generate inspection tasks on the ground, drive the UAV to complete flight shooting and video acquisition preprocessing, and form data to be analyzed; The low-rank decomposition module is used to perform low-rank matrix binary decomposition on the data to be analyzed and output sparse outlier components. The candidate band generation module is used to establish a pixel displacement vector field based on sparse anomaly components and calculate the divergence distribution, locate the anomaly candidate center and generate a ring-shaped candidate band. An improved segmentation module is used to construct an improved PSENet network and perform continuous range field prediction, two-dimensional vector flow prediction and cross-frame potential matching on the annular candidate band, outputting defect instance characterization data. The structured generation module is used to map defect instance characterization data to a three-dimensional track coordinate system, construct dynamic surface deduction, and generate structured inspection results.
[0029] The hierarchical feedback module is used to hierarchically encapsulate and schedule the feedback of the structured inspection results, generate inspection reports, and complete archiving.
[0030] Example 1: To verify the feasibility of this invention in practice, it was applied to a continuous UAV inspection cycle. The UAV payload collected inspection videos and extracted frames on the airborne end to form an online processing frame set. 1800 frames were extracted, with a single frame resolution of 1920×1080 pixels and an extraction frequency of 5 frames per second. The inspected objects exhibited three types of defects: fine cracks, edge gaps, and foreign object attachments. Crack widths ranged from 1 to 5 pixels. Approximately 34% of the cracks were interrupted by strong reflections, and approximately 22% of the defects were located in the periodic stripe area generated by the rotor shadow. Communication link fluctuations were significant, with uplink bandwidth varying from 1.2 to 5.6 megabits per second, round-trip latency varying from 120 to 420 milliseconds, and packet loss rate varying from 1% to 9%. Traditional methods involve collecting data along a preset flight path, transmitting the entire dataset back, offline segmentation at the ground end, and manual verification. This often results in slow alarm arrival, insufficient evidence angles leading to re-flying for re-shooting, and lengthy report processing times.
[0031] After the method of this invention starts running, the UAV preprocesses the extracted frames, scaling the pixel values from 0–255 to 0–1, using a 5×5 median filter to remove isolated noise, and geometrically aligning adjacent frames to reduce jitter. The average signal-to-noise ratio before preprocessing is 16.1 dB, which is improved to 20.4 dB after preprocessing. In the rotor shadow area, the average stripe amplitude decreases from 0.18 to 0.11. A continuous frame observation matrix is constructed by taking frames from smallest to largest timestamp. Each frame is expanded into a pixel vector in a row-first manner and then concatenated column-wise. The observation matrix is divided into 16×16 pixel sub-blocks to form a sub-block array, with 120 horizontal sub-blocks and 67 vertical sub-blocks, for a total of 8040 sub-blocks. Boundaries less than 16 pixels are filled with pixel values of 0.
[0032] Entering the low-rank matrix binary decomposition stage, block cyclic low-rank convolution decomposition is first performed. The sub-block array is organized with 67 rows and 120 sub-blocks per row. The second row is cyclically shifted one sub-block position to the right relative to the first row, and the third row is cyclically shifted one sub-block position to the right relative to the second row, and so on, to obtain the block cyclic matrix. Alternating least squares is used to update the low-rank factor and low-rank coefficient in the block cyclic domain, with 20 iterations, to obtain the low-rank background component. This is then subtracted from the observation matrix to obtain the residual matrix. At this stage, the root mean square error of low-rank reconstruction has decreased from 0.142 to 0.031. Next, complex-domain phase-sharing decomposition is performed, mapping the residual matrix to a complex-domain residual matrix. The complex elements consist of amplitude and phase, with amplitude representing the residual brightness structure and phase representing periodic perturbations. After 20 iterations in the complex domain, the complex-domain low-rank periodic component is obtained. The complex-domain sparse anomaly component is obtained by subtracting the complex-domain low-rank periodic component from the complex-domain residual matrix, and the amplitude is taken to obtain the sparse anomaly component. Statistics show that the proportion of pixels with an amplitude greater than 0.2 in the sparse outlier components is 0.46%, while the proportion in the periodic stripe region has decreased from 0.71% without phase-sharing decomposition to 0.39%, indicating that the periodic perturbation has been effectively removed.
[0033] Sparse outlier components are used to generate annular candidate bands. For each frame's 2D pixel array, a 3×3 Sobel operator is used to calculate the horizontal and vertical gradients, forming a pixel displacement vector field. Then, a divergence map is calculated using central difference and normalized to 0–1. Local maxima in the divergence map are identified as outlier candidate centers. A closed trajectory is formed by stepping along the divergence gradient, and the region inside the closed trajectory is merged with its neighboring region to obtain the annular candidate band. This cycle generates 156 annular candidate bands, with each band covering an average of 14,800 pixels, representing approximately 0.71% of the entire frame. By limiting the segmentation calculation to the effective region, the inference time per frame is reduced from 92 milliseconds for full-frame processing to 37 milliseconds.
[0034] An improved PSENet network with a circular candidate band input is implemented. The implicit range field output branch directly outputs a continuous range field and generates a range level set based on a distance constant. The decision criterion is that the absolute value of the difference between the normalized range value and the distance constant is no greater than 0.02, forming a kernel set. The 2D vector flow output branch outputs a 2D vector flow, with vector amplitudes truncated to an upper limit of 1.0 pixels. Each pixel is iteratively updated along the vector flow to be assigned to a kernel, with an upper limit of 30 iterations, resulting in a defect instance mask. Cross-frame gravitational potential energy mapping calculates the potential energy contribution based on the instance centroid, achieving minimum potential energy matching to maintain consistent instance identifiers. This period outputs 92 defect instance characterization data: 58 cracks, 21 notches, and 13 foreign objects. The cross-frame instance identifier switching count is 3, with an identifier stability rate of 96.7%, compared to 81.4% for traditional frame-by-frame offline segmentation.
[0035] The comparative experiments used the same batch of data. The traditional method involved full data transmission + offline segmentation network + manual verification, while the present invention used airborne low-rank decomposition + candidate bands + improved PSENet + hierarchical transmission. Both methods used 4000 training samples and 1000 test samples. The results showed that the defect pixel segmentation accuracy was 82.6% for the traditional method and 91.3% for the present invention; the mean cross-intersection over union (CUI) ratio of defect instances was 0.71 for the traditional method and 0.83 for the present invention; the fine crack recognition rate was 54.8% for the traditional method and 79.6% for the present invention; the crack breakage rate was 26.4% for the traditional method and 10.2% for the present invention; and the cross-frame instance identification stability rate was 81.4% for the traditional method and 96.7% for the present invention. Under link fluctuation conditions, the traditional method had a 168-second delay in the arrival of the first alarm and unstable arrival of key evidence; the present invention had a 8–19-second delay in the arrival of alarm summaries after hierarchical transmission, a 31–74-second delay in the arrival of evidence clipping, and the original data was retransmitted during the link recovery phase without loss.
[0036] Regarding rework and re-flight procedures, the traditional method resulted in a 33% re-flight reshoot rate, while this invention reduced it to 9% due to online evidence collection and linkage. Report compilation and archiving time decreased from 36 minutes to 9 minutes. The proportion of samples awaiting review decreased from approximately 18% under traditional manual screening to 9.3% after automatic verification using this invention. Data shows that this invention can simultaneously improve defect detection accuracy, cross-frame stability, alarm timeliness, and evidence archiving efficiency in scenarios with weak networks and periodic disturbances.
[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A UAV inspection method based on air-ground cooperation and AI visual detection, characterized in that, include: Inspection tasks are generated on the ground and sent to the drone. The drone performs flight and shooting according to the inspection tasks and collects video data. The video data is preprocessed to form data to be analyzed. Perform low-rank matrix binary decomposition, rearrange the data to be analyzed into a block cyclic matrix based on the block cyclic low-rank convolution solution, complete the calculation of low-rank components in the block cyclic domain, and use complex domain phase-shared decomposition to represent the data to be analyzed as complex domain factors, and map the amplitude components and phase components into brightness structure and periodic perturbation respectively to obtain sparse anomaly components. A pixel displacement vector field is established based on sparse anomaly components. The divergence distribution of the pixel displacement vector field is calculated. Anomaly candidate centers are marked at local maxima of divergence. A ring-shaped candidate band is generated by continuously tracking along the divergence gradient. An improved PSENet network is constructed to perform defect segmentation and instance attribution processing on the annular candidate band. Implicit distance field kernels are introduced to make the continuous distance field prediction explicit. Two-dimensional vector flow prediction is performed based on vector flow homotopy attribution. Pixels are merged into the corresponding kernels along the vector flow trajectory. Potential energy matching is performed using cross-frame gravitational potential energy mapping. Defect instance characterization data is output. The defect instance characterization data is mapped to the current three-dimensional trajectory coordinate system of the UAV, a dynamic surface is constructed, and continuous deduction is performed in the tangent direction of the dynamic surface to generate the structured inspection results. Based on link status parameters and backhaul priority parameters, the inspection structured structure is hierarchically encapsulated and backhaul scheduled to generate an inspection report and complete archiving. 2.The UAV inspection method based on air-ground cooperation and AI visual detection of claim 1, wherein, The inspection task includes a sequence of waypoints arranged in chronological order, the flight altitude, flight speed, heading angle range, gimbal pitch angle range, zoom range, and shooting time markers for each waypoint. 3.The UAV inspection method based on air-ground cooperation and AI visual detection of claim 1, wherein, The preprocessing of video data to form data to be analyzed includes: On the ground, inspection tasks are generated based on the geographical range of the target to be inspected, the boundary of the no-fly zone, the coordinates of the take-off and landing points, and the operation time window. The inspection tasks are then encoded into task data packets. The mission data packet is sent to the UAV terminal through the air-to-ground communication link. The UAV terminal performs integrity verification and parsing of the mission data packet, establishes the mapping relationship between mission identifier and waypoint sequence, writes the waypoint sequence into the flight control queue, and writes the shooting time identifier and gimbal control parameters into the payload control queue. The UAV arrives at each waypoint in sequence according to the flight control queue and performs shooting according to the payload control queue. It collects video data and synchronously records the timestamp, positioning coordinates, attitude angle and gimbal parameters corresponding to each frame of video data. The video data is preprocessed to form data to be analyzed. The preprocessing includes frame extraction at a fixed frame rate, resolution unification of the extracted frames, geometric alignment of the extracted frames, and assembling the extracted frames into time-series data segments according to the number of consecutive frames.
4. The unmanned aerial vehicle inspection method based on air-ground cooperation and AI visual detection according to claim 1, characterized in that, The obtained sparse outlier components include: The data to be analyzed is constructed into an observation matrix in the order of consecutive frames. The observation matrix is formed by stitching multiple frames of images together column by column. Each column is a pixel vector obtained by expanding a frame of images in row priority order. The observation matrix is divided into sub-block arrays according to the sub-block size. Based on the block-circular low-rank convolution solution, the sub-block array is rearranged in a block-circular manner to form a block-circular matrix. The block-circular matrix satisfies that the sub-blocks in the next row are cyclically shifted relative to the sub-blocks in the previous row by a fixed number of sub-blocks. The low-rank factor is iteratively updated in the block-circular domain to obtain the low-rank factor matrix and the low-rank coefficient matrix. The low-rank background component is obtained by multiplying the transpose of the low-rank factor matrix and the low-rank coefficient matrix. The residual matrix is obtained by subtracting the observation matrix from the low-rank background component. The residual matrix is represented as a complex domain residual matrix by using complex domain phase-sharing decomposition. The complex domain residual matrix is decomposed into a complex domain low-rank periodic component and a complex domain sparse anomaly component. The complex domain low-rank periodic component is obtained by multiplying the complex domain basis matrix and the transpose of the complex domain coefficient matrix. Each element of the complex domain basis matrix and the complex domain coefficient matrix consists of an amplitude component and a phase component. The amplitude component represents the brightness structure in the residual, and the phase component represents the periodic perturbation in the residual. The sparse outlier component in the complex field is obtained by the difference between the observation matrix and the low-rank background component in the complex field. The absolute value operation of the amplitude of the sparse outlier component in the complex field is performed to obtain the sparse outlier component. The amplitude of each element in the sparse outlier component is equal to the square root of the sum of the squares of the real part and the squares of the imaginary part of the element.
5. The unmanned aerial vehicle inspection method based on air-ground cooperation and AI visual detection according to claim 1, characterized in that, The generation of the annular candidate band includes: Using the sparse outlier component corresponding to each frame of two-dimensional pixel array as input, spatial gradient calculation is performed on the two-dimensional pixel array to obtain the horizontal gradient field and the vertical gradient field. The horizontal gradient field and the vertical gradient field at each pixel position form the pixel displacement vector field. The vector of the pixel displacement vector field at each pixel position is composed of the horizontal component and the vertical component. The divergence distribution is calculated for the pixel displacement vector field. The divergence value at any pixel position is equal to the sum of the partial derivatives of the horizontal component of the pixel position with respect to the horizontal coordinate and the partial derivatives of the vertical component with respect to the vertical coordinate, forming a divergence map. Local maxima in the divergence map are identified and used as candidate anomaly centers. A closed trajectory is generated by stepping outward from the candidate anomaly centers along the gradient direction of the divergence map. The internal region of the closed trajectory is merged with the neighboring region of the closed trajectory to form an annular candidate band.
6. The unmanned aerial vehicle inspection method based on air-ground cooperation and AI visual detection according to claim 1, characterized in that, The output defect instance characterization data includes: An improved PSENet network is constructed and the pixel features corresponding to the annular candidate band are input. The improved PSENet network sets an implicit range field output branch at the segmentation output position of the original PSENet network to make the implicit range field kernel explicit. A two-dimensional vector flow is set at the pixel-level prediction output position of the original PSENet network to perform vector flow homotopy assignment. After the instance generation output, cross-frame gravitational potential energy mapping is performed at the cross-frame identifier mapping position. The continuous range field obtained from the implicit range field output branch is kernel-explicitly processed based on the implicit range field kernel explicitization, generating multiple sets of range level sets and generating a kernel set from the pixel set corresponding to each range level set. The range level set is the set of pixels that satisfy the distance value equals the distance constant. Based on the homotopy assignment of vector flow, the two-dimensional vector flow obtained from the output branch of the two-dimensional vector flow is processed for instance assignment. Starting from each pixel, the two-dimensional vector flow is iteratively updated along the two-dimensional vector flow to obtain the pixel trajectory. The pixel coordinates are updated to the sum of the two-dimensional vectors at the previous updated pixel coordinates and the current pixel coordinates until the pixel trajectory enters any kernel region in the kernel set. The kernel identifier corresponding to the entered kernel region is assigned as the defect instance identifier of the pixel to generate the defect instance mask of the current frame. Cross-frame gravity potential energy mapping is used to perform cross-frame consistency processing on the defect instance identifier of the current frame. The potential energy contribution is calculated at the centroid position of each defect instance mask in the previous frame and the potential energy value is calculated at the centroid position of the defect instance in the current frame. The defect instance identifier of the previous frame corresponding to the minimum potential energy value is assigned as the defect instance identifier of the current frame, and the defect instance characterization data is output. The improved PSENet network is trained by constructing training samples and updating network parameters. The training samples include circular candidate band images and corresponding annotations. The network parameters are updated using the total loss, which is equal to the sum of the segmentation loss corresponding to the pixel annotation of the defect region, the distance field loss corresponding to the continuous distance field annotation, the vector flow loss corresponding to the two-dimensional vector flow annotation, and the consistency loss corresponding to the defect instance identifier annotation of the adjacent frame.
7. The UAV inspection method based on ground-air collaboration and AI visual detection according to claim 1, characterized in that, The generation of structured inspection results includes: Each defect instance mask in the defect instance characterization data is converted into a set of pixel coordinates. The centroid coordinates of the pixel coordinate set are calculated and the boundary coordinate set of the defect instance is determined. The geometric quantization parameters of the defect instance are calculated based on the boundary coordinate set. Based on timestamps, positioning data, pose data and gimbal parameters, establish a mapping relationship from pixel coordinates to a three-dimensional track coordinate system, map the centroid coordinates and boundary coordinates set to the defect center coordinates and defect boundary coordinates set, and associate defect instance identifiers and evidence indexes to the defect center coordinates and defect boundary coordinates set. A dynamic surface is constructed based on the set of defect center coordinates and defect boundary coordinates. The dynamic surface is defined by the current position of the UAV, the range of allowable flight altitude, the range of allowable pitch angle, and the range of safe distance. The reshoot state sequence is continuously generated in the tangent direction of the dynamic surface according to the step size. The reshoot state sequence and defect instance characterization data are jointly encapsulated to generate the inspection structured result.
8. The UAV inspection method based on ground-air collaboration and AI visual detection according to claim 1, characterized in that, The process of hierarchical encapsulation and backhaul scheduling, generating inspection reports and completing archiving includes: Collect link status parameters of the air-to-ground communication link and generate link status identifiers. The link status parameters include uplink bandwidth, downlink bandwidth, round-trip time, and packet loss rate. The inspection structured results are divided into alarm summary data, evidence trimming data and raw data according to data type, and data packages are generated for each. A backhaul queue is established based on the backhaul priority parameter and scheduled in conjunction with the link status identifier. The backhaul queue is arranged in the order of alarm summary data, evidence clipping data, and original data. Data packets are backhauled sequentially within the allowed backhaul range corresponding to the link status identifier, and data packets exceeding the allowed backhaul range are buffered. When the allowed backhaul range corresponding to the link status identifier expands, the buffered data packets are supplemented according to the backhaul queue order. The ground end verifies and unpacks the received data packets and generates an inspection report for archiving.
9. A drone inspection system based on ground-air collaboration and AI visual detection, comprising executing the drone inspection method based on ground-air collaboration and AI visual detection as described in any one of claims 1 to 8, characterized in that, Includes the following modules: The task assignment module is used to generate inspection tasks on the ground, drive the UAV to complete flight shooting and video acquisition preprocessing, and form data to be analyzed; The low-rank decomposition module is used to perform low-rank matrix binary decomposition on the data to be analyzed and output sparse outlier components. The candidate band generation module is used to establish a pixel displacement vector field based on sparse anomaly components and calculate the divergence distribution, locate the anomaly candidate center and generate a ring-shaped candidate band. An improved segmentation module is used to construct an improved PSENet network and perform continuous range field prediction, two-dimensional vector flow prediction and cross-frame potential matching on the annular candidate band, outputting defect instance characterization data. The structured generation module is used to map defect instance characterization data to a three-dimensional track coordinate system, construct dynamic surface deduction, and generate structured inspection results.
10. The hierarchical feedback module is used to hierarchically encapsulate and schedule the feedback of the structured inspection results, generate inspection reports, and complete archiving.