Unmanned aerial vehicle-mounted visual online detection and target identification system
By using gradient analysis and feature verification modules, the high computational complexity of UAV visual inspection systems under image degradation conditions is solved, achieving low-cost and stable target recognition results.
Patent Information
- Application Number
- CN202511047877.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing UAV-borne visual inspection systems rely on computationally complex deep learning models when faced with image degradation, resulting in excessive computing power and power consumption, making it difficult to achieve stable target recognition on resource-constrained UAV platforms.
By employing a gradient analysis module, an image validity diagnosis module, and a feature credibility verification module, the statistical distribution of image gradient directions is simplified and quantized to directly capture target structural information, thereby reducing computational costs and improving recognition stability.
Under the condition of limited UAV resources, target recognition with low computational cost was achieved, which improved the stability and recognition reliability of the system in complex environments and avoided unnecessary computational overhead and misleading conclusions.
Smart Images

Figure CN120953840A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an unmanned aerial vehicle (UAV)-borne visual online detection and target recognition system, belonging to the field of computer vision and image recognition technology. Background Technology
[0002] Currently, in the application of UAV-borne visual inspection, the industry has widely recognized the use of deep learning-based image recognition systems to achieve high-precision automated detection of remote targets. These systems typically follow a technical path: first, image enhancement algorithms are used to improve the pixel-level quality of the image to be analyzed; then, a complex deep neural network model is used to extract high-dimensional features; and finally, the target is classified and identified. Under ideal imaging conditions, this technical path can indeed achieve satisfactory detection accuracy and constitutes the current technological foundation of this field.
[0003] However, when this mainstream technology path is pushed to large-scale, long-endurance, all-weather autonomous inspection applications, the cost of maintaining its accuracy advantage becomes increasingly apparent. As an edge computing platform with strict constraints on power consumption, computing power, and load, drones face unavoidable image degradation issues in the real world, such as motion blur, low light, or haze. To compensate for information loss, existing technologies are forced to continuously upgrade the complexity of their feature extraction models, such as introducing computationally intensive self-attention mechanisms or deformable convolutional networks. This trend of continuously increasing computing power requirements in pursuit of accuracy contradicts the engineering goals of lightweight, low-cost, and long-endurance drone platforms. To resolve this contradiction, a seemingly direct improvement approach is to equip drones with onboard processing units. However, this approach actually introduces new performance constraints, because a stronger processor means a larger size, heavier load, and higher power consumption. This will directly reduce the effective operating radius and endurance of drones and increase the deployment cost of a single mission, making automated inspection, which was originally intended to improve economic efficiency, impractical due to the rising hardware costs.
[0004] Specifically, existing technologies suffer from the following shortcomings: 1. The system is highly dependent on the pixel-level quality of images. Once the image quality is severely degraded due to environmental factors, the effectiveness of the entire recognition chain becomes unsustainable. 2. The complex models used to handle images result in huge computational and power consumption overhead for the airborne system, conflicting with the resource limitations of the UAV platform and restricting the universal application of the technology. 3. There is a lack of a recognition method that can extract robust structured features at a low computational cost and make reliable decisions even when image information is incomplete. Therefore, how to establish a recognition method that is free from dependence on high-quality images and complex models, and that can directly extract stable structural information of the target from degraded images at a low computational cost and make reliable recognition under the condition of limited UAV resources, is the technical problem to be solved by this invention. Summary of the Invention
[0005] This invention provides an unmanned aerial vehicle (UAV)-borne visual online detection and target recognition system. Its main purpose is to solve the problem of how to reliably recognize targets from degraded images with low computational cost under the condition of limited UAV resources.
[0006] To achieve the above objectives, the present invention provides an unmanned aerial vehicle (UAV)-borne visual online detection and target recognition system, the system comprising:
[0007] A gradient analysis module is configured to divide an image acquired by an airborne image sensor into multiple image blocks, and for each image block, generate a gradient magnitude statistical distribution and a spatial orientation map composed of the dominant gradient directions of each image block based on pixel gradient calculation.
[0008] An image validity diagnosis module, whose input is connected to a gradient analysis module, is configured to compare the proportion of high-amplitude gradient pixels in the gradient magnitude statistical distribution with a sharpness threshold stored in an onboard non-volatile memory, and generate an image validity signal characterizing whether the current image contains valid information based on the comparison result.
[0009] A feature confidence verification module, whose input is connected to the gradient analysis module, is configured to generate a feature confidence weight that characterizes the authenticity of gradient information based on the analysis of the spatial relationship continuity between the dominant gradient directions of adjacent image blocks in the spatial orientation map.
[0010] A target decision module, whose input is connected to a gradient analysis module, an image validity diagnosis module, and a feature credibility verification module, is configured to only combine the pixel clustering degree and feature credibility weight of each image patch with a gradient-target mapping table that is also pre-stored in an onboard non-volatile memory when an image validity signal is received indicating that the image contains valid information, and then match and determine the target type of the image patch from a gradient-target mapping table that is also pre-stored in an onboard non-volatile memory.
[0011] Preferably, the gradient analysis module is further configured to: determine the dominant gradient direction and pixel clustering by performing voting statistics on the pixel gradient directions in four discrete directions within the image block. The pixel clustering is the ratio of the number of votes in the dominant gradient direction to the total number of votes, thus avoiding the need for precise trigonometric function calculations of the gradient angle.
[0012] Preferably, the image validity diagnosis module is further configured to: when the proportion of high-amplitude gradient pixels is determined to be lower than the sharpness threshold, generate an image validity signal indicating that the image does not contain valid information, and the target decision module will stop running after receiving the signal to avoid the energy waste caused by invalid calculations when the sensor field of view is obstructed.
[0013] Preferably, the feature confidence verification module is further configured to: calculate the continuity entropy H of the spatial orientation map based on a set of transition probabilities p(i) between adjacent image blocks of the dominant gradient direction, where p(i) is the probability of transitioning from one dominant gradient direction to another; and generate feature confidence weights based on the calculation result of the continuity entropy H, wherein the value of the continuity entropy H is negatively correlated with the value of the feature confidence weights.
[0014] Preferably, the target decision module is further configured to: multiply the pixel clustering degree with the feature confidence weight to obtain a weighted clustering degree, and compare the weighted clustering degree with a decision threshold stored in an onboard non-volatile memory, using the comparison result as the basis for matching from the gradient-target mapping table.
[0015] Preferably, the gradient-target mapping table contains multiple entries, each of which defines a unique target type and a set of corresponding decision rules. The decision rules include a dominant gradient direction type and a weighted clustering range.
[0016] Preferably, the gradient analysis module uses the Sobel operator to calculate pixel gradients; and the four discrete directions include the horizontal direction, the vertical direction, and two diagonal directions.
[0017] Preferably, the target decision module is further configured to: when the weighted clustering degree is not higher than the decision threshold, the image block is determined to be a background region, and no subsequent target type matching and output are performed, so as to reduce the overall computational load of the system.
[0018] Preferably, the proportion of high-amplitude gradient pixels is obtained by counting the number of pixels with gradient amplitudes greater than the sharpness threshold and then dividing by the total number of pixels.
[0019] Preferably, the system is physically manifested as an embedded computing unit integrated into the UAV platform. The logical functions of the gradient analysis module, image validity diagnosis module, feature credibility verification module, and target decision module are all executed by one or more microcontroller units configured with firmware. The microcontroller units are connected to the airborne image sensor and the airborne non-volatile memory.
[0020] Compared with the prior art, the beneficial effects of the present invention are:
[0021] 1. This invention establishes a recognition method that directly makes decisions based on the structural features of images. This method decouples the recognition process from the pixel-level quality of the image. When UAVs face image degradation such as motion blur or low light, the system no longer relies on computationally expensive image enhancement or reconstruction processes. Instead, it directly captures the structural information of the target that remains stable despite the image degradation process by performing minimal quantization of the statistical distribution of the image gradient direction. Thus, on resource-constrained airborne platforms, it achieves effective recognition of predetermined targets with lower computing power consumption and energy overhead. This provides a technical solution for the application of UAV visual inspection functions in terms of cost and power consumption.
[0022] 2. This invention generates gradient amplitude statistical distribution and spatial orientation map in parallel through a gradient analysis module, and performs pre-analysis on the former using an image validity diagnosis module, thereby giving the system a sensory self-diagnosis capability before performing core recognition tasks. When the UAV encounters scenarios where the sensor field of view is severely obstructed, such as dense fog or sandstorms, the system can pre-determine whether the current image frame has degenerated into an invalid frame without effective information based on the attenuation of gradient amplitude in the global range, and stop the subsequent recognition process accordingly. This design avoids the system from continuously performing futile analysis and calculation on meaningless data when the sensor field of view is severely obstructed, which not only saves valuable airborne energy, but also prevents the system from outputting misleading conclusions due to the failure of the input source, thus improving the safety of the overall decision-making.
[0023] 3. By introducing a feature credibility verification module, this invention adds a verification step regarding the authenticity of feature sources to the recognition decision-making process after image validity diagnosis. This module quantifies the spatial continuity of the gradient field by analyzing the transfer relationship of the dominant gradient direction between adjacent image blocks in the spatial orientation map, thereby distinguishing whether the low gradient clustering in a region originates from a featureless physical background or from structural damage caused by transient optical noise. This mechanism enables the system to identify and suppress or mark the pseudo-gradients generated by complex electromagnetic and optical interferences such as arc discharge or high-frequency reflection, thereby avoiding the core recognition logic being misled by false signals and ensuring the reliability of recognition results in complex interference environments.
[0024] 4. This invention constructs a logically progressive collaborative decision-making system consisting of three modules: image validity diagnosis, feature credibility verification, and target decision-making. The system first confirms whether there is sufficient valid information in the image for analysis, then verifies whether the extracted structural features originate from the real physical world rather than noise interference, and finally, under the dual premise of ensuring the validity and credibility of the information, performs the final target matching and recognition. This interconnected internal verification process integrates three originally isolated analysis dimensions into an organic whole with multi-level and endogenous robust capabilities. This improves the overall stability and environmental adaptability of the system when dealing with the ever-changing, complex, and uncertain visual detection tasks in the real world, compared to existing technologies that only optimize a single recognition algorithm. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the functional modules and data flow of the visual online detection and target recognition system of the present invention;
[0026] Figure 2 This is a comparison chart of the detection rate performance of the system of this invention and the traditional deep learning system under different environmental conditions;
[0027] Figure 3 This is a schematic diagram of the logic flow and data interaction of the visual online detection and target recognition system of the present invention.
[0028] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in further detail below. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0030] This invention provides an unmanned aerial vehicle (UAV)-borne visual online detection and target recognition system. Physically, this system can be embodied as an embedded computing unit integrated into the UAV platform. Its logical functions are executed by one or more microcontroller units configured with firmware, such as ARM Cortex-M7 series microcontrollers. The system includes a gradient analysis module, an image validity diagnosis module, a feature confidence verification module, and a target decision module. The output of the gradient analysis module is connected to the inputs of the other three modules, while the outputs of the image validity diagnosis module and the feature confidence verification module are both connected to the input of the target decision module. This constitutes a system that executes sequentially... The system employs a decision-making process involving effectiveness diagnosis, credibility verification, and target matching. In a specific application scenario, such as when a drone autonomously inspects power lines to detect insulator cracks or foreign objects in conductors in real time, the gradient analysis module is configured to execute an image structure information extraction process suitable for efficient operation on a microcontroller to address image quality degradation caused by drone vibration or changes in ambient lighting. After this process is initiated, the gradient analysis module divides an image frame acquired by an onboard image sensor into multiple N×N image blocks, where N can be set to 16. Then, for each image block, the module applies a Sobel operator to calculate the x-direction gradient G of the pixels within the block. x and gradient G in the y direction y To avoid performing square root and inverse tangent calculations on the microcontroller, this module does not calculate precise gradient magnitude and angle. Instead, it determines the pixel clustering and dominant gradient direction of each image block by performing a voting statistical analysis of the pixel gradient direction in four discrete directions: horizontal, vertical, and two diagonal directions. This voting statistical analysis compares G... x With G y The sign and absolute value of the gradient are implemented entirely through integer operations; after the statistics are completed, the pixel clustering degree of an image patch is determined as the ratio of the number of votes in the dominant gradient direction to the total number of votes; at the same time, this module calculates an approximate gradient magnitude in parallel, such as calculating |G x |+|G y Based on this, a gradient magnitude statistical distribution is generated; finally, the gradient analysis module outputs the gradient magnitude statistical distribution and a spatial orientation map composed of the dominant gradient directions of each image patch, and transmits the pixel clustering degree and dominant gradient direction information of each image patch to the target decision module, providing a structured data foundation for subsequent decision-making under image degradation conditions.
[0031] Given that drones may encounter dense fog or sandstorms during flight, resulting in significant loss of image information, to avoid the system performing calculations on data without analytical value under such circumstances, the image validity diagnosis module in the system is configured to execute an image validity diagnosis procedure before the core recognition task. In this procedure, the module receives the gradient amplitude statistical distribution from the gradient analysis module and calculates the proportion of high-amplitude gradient pixels based on this. This proportion is calculated by counting the number of pixels in the image whose gradient amplitude is greater than a preset sharpness threshold, and then dividing that number by the total number of pixels. The sharpness threshold is set through a deterministic offline calibration process. This process includes acquiring a set of baseline images under clear weather conditions and a set under dense fog conditions where visibility does not meet mission requirements. The distribution of the proportion of high-amplitude gradient pixels is calculated for each set, and the value that can effectively distinguish the two distributions is selected as the sharpness threshold. This value is then stored in onboard non-volatile memory. The module compares the real-time calculated proportion with the stored sharpness threshold. If the proportion is lower than the threshold, an image validity signal indicating that the current image does not contain valid information is generated and transmitted to the target decision module. Upon receiving this signal, the decision-making module suspends subsequent recognition steps for the current image frame. Even if the image passes the validity check, its gradient information may be affected by random optical noise interference caused by factors such as high-voltage arcs or metallic reflections. This interference can create spatially discontinuous pseudo-gradients, affecting the recognition results. Therefore, the feature credibility verification module is configured to execute a feature source credibility verification procedure after the image validity check. This module receives the spatial orientation map generated by the gradient analysis module, and its judgment is based on the fact that the structural gradient of a real object has spatial continuity, while the light... The gradient of the noise tends to be random in space. To quantify this spatial continuity, this module calculates the continuity entropy of the spatial orientation map H = ∑p(i)log2(p(i)) based on a set of transition probabilities p(i) between adjacent image patches of the dominant gradient direction, where p(i) is the probability of transitioning from one dominant gradient direction to another. The log2 operation can be implemented on the microcontroller using a lookup table. The calculated continuity entropy H is negatively correlated with the feature confidence. Based on this, the module generates a feature confidence weight W based on the calculation result of H, for example, through W = 1 - H / H max The mapping relationship is generated and output to the target decision module;
[0032] Ultimately, the target decision module is activated only upon receiving an image validity signal indicating that the image contains valid information. After activation, this module combines the pixel clustering degree of each image patch obtained from the gradient analysis module with the feature confidence weights obtained from the feature confidence verification module to perform a recognition decision. Specifically, the module multiplies the pixel clustering degree by the feature confidence weights to obtain a weighted clustering degree, and compares this weighted clustering degree with a decision threshold pre-stored in onboard non-volatile memory. If the weighted clustering degree is not higher than the decision threshold, the image patch is determined to be a background region, and no further matching is performed. Conversely, if the weighted clustering degree is higher than the decision threshold, the system uses the dominant gradient direction of the image patch as a reference to identify the target region. A gradient-target mapping table pre-installed in onboard non-volatile memory is used to match and determine the target type of image patches. This gradient-target mapping table contains multiple entries, each defining the association between a target type and a set of corresponding decision rules. The decision rules include the type of a dominant gradient direction and a weighted clustering range. For example, an entry can be defined as: when the dominant gradient direction is vertical and the weighted clustering is in the range of 0.9 to 1.0, the target type is determined to be an insulator crack. Because this decision method directly relies on the statistical characteristics of gradient information rather than the pixel-level quality of the image, it can reduce the computing power consumption of the onboard system while dealing with image degradation of specific types, thereby completing the recognition task.
[0033] Example 1: In an autonomous inspection mission of a power transmission line, a drone equipped with the visual online detection and target recognition system of this invention enters a mountainous section of the line for operation in the early morning. During this period, two image degradation conditions occur simultaneously: fog in some areas reduces image contrast, and sunlight forms intermittent, high-intensity specular reflections on the wet surfaces of insulators and conductors. When the drone flies into an area with dense fog, the image frames acquired by its onboard image sensor become blurred overall, and the edge contour information of objects is weakened. The gradient analysis module in the system processes the image as described in the aforementioned specific implementation, and the resulting gradient amplitude statistical distribution shows that the proportion of high-amplitude gradient pixels is low. After receiving this distribution, the image validity diagnosis module calculates that the proportion of high-amplitude gradient pixels is lower than the preset sharpness threshold in the onboard non-volatile memory. Therefore, the module generates and sends an image validity signal to the target decision module indicating that the current image does not contain valid information. Based on this signal, the target decision module stops all subsequent processing of the image frame.
[0034] Subsequently, the drone flew out of the fog area and entered a region with complex lighting. A series of ceramic insulators simultaneously exhibited a longitudinal micro-crack caused by aging and multiple randomly distributed high-intensity specular reflection points caused by sunlight. Under these conditions, the image validity diagnosis module determined that the proportion of high-amplitude gradient pixels in the current frame was higher than the sharpness threshold, thus classifying the image as valid and continuing the recognition process. The spatial orientation map generated by the gradient analysis module showed that, within the neighborhood of the image patch containing the real crack, the dominant gradient direction exhibited spatial continuity in the vertical direction, while in the region containing the reflection points, the dominant gradient direction showed a random spatial distribution. After receiving this spatial orientation map, the feature credibility verification module calculated the continuity entropy H for the two regions. The H value for the real crack region approached zero, thus obtaining a feature credibility weight W close to 1, while the H value for the reflection point region was higher, obtaining a feature credibility weight W close to 0. Finally, in the target decision module, for... The image patch containing the crack has a pixel clustering degree multiplied by a feature confidence weight W close to 1. The resulting weighted clustering degree is higher than the decision threshold. The system then matches and outputs the target type and location of the insulator crack from the gradient-target mapping table based on its vertical dominant gradient direction. For the image patch containing the reflection point, after multiplying by a feature confidence weight W close to 0, the weighted clustering degree is lower than the decision threshold. Therefore, it is judged as a background area, and its reflected light point information does not interfere with the final result. The entire recognition process does not require image enhancement or rely on a computationally intensive neural network model. It only requires double pre-verification and statistical decision-making of gradient information and runs on a single low-power microcontroller, solving the constraint between detection performance and platform resource consumption in visual inspection applications. The inspection log output by the system finally records target information that has passed internal verification of validity and confidence, forming a recognition result that does not depend on ideal imaging conditions.
[0035] Example 2: To compare the performance of the system of the present invention with a control group recognition system based on a deep neural network under different image qualities, an indoor hardware-in-the-loop simulation test platform was built. The platform includes an industrial camera fixed to the end of a six-axis robotic arm to simulate the flight attitude and jitter of a drone, an insulator with a longitudinal crack of known size placed within the camera's field of view as the target, and a programmable environmental simulation device that can generate uniform fog through an aerosol generator and transient optical noise through an LED array. During the experiment, the video stream output by the camera was simultaneously input to two independent computing units. One was an ARM Cortex-M7 microcontroller unit equipped with the system of the present invention, forming the experimental group, and the other was an embedded computing unit running an optimized target recognition model, forming the control group.
[0036] Before the experiment, key parameters of the experimental system, such as the sharpness threshold and decision threshold, need to be calibrated. Taking the sharpness threshold as an example, its setting needs to balance the ability to perceive effective information under low contrast and the ability to filter invalid and blurred images. The calibration procedure is as follows: First, a set of images taken under clear lighting and containing effective target outlines are acquired, along with another set of invalid images taken in dense fog where target outlines cannot be identified. Then, for each image in the two image sample sets, the proportion of high-amplitude gradient pixels is calculated through the gradient analysis module. Finally, the statistical distribution of the two proportion data is analyzed, and a value that can effectively distinguish the two distributions is selected as the sharpness threshold and stored in the onboard non-volatile memory. During the experiment, the robotic arm and the environmental simulation equipment generate five test conditions in sequence according to the preset script: clear condition, motion blur, low light, medium density fog, and clear condition with added transient optical noise. Under each condition, the system continuously processes 3000 frames of images and records the target detection rate, false alarm rate, average processor load, and average power consumption of the computing unit for the experimental and control groups. Table 1 shows the data recorded in the experiment.
[0037] Table 1: Comparison of performance data between the experimental group and the control group under different operating conditions.
[0038] Test conditions Test Project experimental group control group Clear conditions Detection rate (%) 98.2 99.5 Processor load (%) 12.5 85.4 Average power consumption (W) 0.8 7.5 Motion blur Detection rate (%) 95.5 65.3 False alarm rate (%) 0.5 2.1 low light Detection rate (%) 94.8 70.1 fog Detection rate (%) 91.3 15.7 Image validity signal trigger rate (%) 8.2 N / A Optical noise Detection rate (%) 96.1 75.2 False alarm rate (%) 0.3 18.6
[0039] Referring to Table 1, under clear conditions, the detection rate of the control group was higher than that of the experimental group, but its processor load and average power consumption were also several times higher. When the image quality degraded due to motion blur, low light, and fog, the detection rate of the control group decreased because its model relied on pixel-level clear features. The detection rate of the experimental group was less affected by image quality degrade because its gradient analysis module focused on processing gradient structure information. At the same time, under fog conditions, its image validity diagnosis module was triggered, actively suspending the calculation of some invalid frames. Under optical noise conditions, the false alarm rate of the control group increased, while the feature credibility verification module of the experimental group suppressed random noise by calculating the continuous entropy H, and its false alarm rate remained at a low level.
[0040] Example 3: This example combines Figures 1 to 3 This describes a UAV-borne visual online detection and target recognition system, such as... Figure 1As shown, an airborne image sensor acquires raw images and transmits them to a gradient analysis module. This module processes the raw images and generates three types of information: gradient magnitude statistical distribution, spatial orientation map, and dominant gradient direction and pixel clustering. The gradient magnitude statistical distribution is sent to an image validity diagnosis module, which determines whether the image contains valid information based on the proportion of pixels with high-amplitude gradients and generates an image validity signal. The spatial orientation map is sent to a feature credibility verification module, which analyzes the authenticity of gradient information based on the continuity of the spatial orientation map and generates a feature credibility weight. Finally, a target decision module receives the dominant gradient direction and pixel clustering from the gradient analysis module, the image validity signal from the image validity diagnosis module, and the feature credibility weight from the feature credibility verification module. It then combines validity, credibility, and pixel clustering for target matching and recognition. If the weighted clustering is not higher than a decision threshold, the corresponding area is classified as a background area, and subsequent matching is stopped. If the weighted clustering is higher than the decision threshold, target type matching and output are performed, such as outputting the identification result of an insulator crack.
[0041] like Figure 2 As shown in the figure, the vertical axis represents the detection rate (%), and the horizontal axis represents five environmental conditions: clear conditions, motion blur, low light, moderate fog, and optical noise. As can be seen from the figure, the solid line, which represents the system of the present invention, has a significantly higher detection rate than the dashed line, which represents the traditional deep learning system, under image degradation conditions such as motion blur, low light, moderate fog, and optical noise. This demonstrates the reliable recognition capability of the system of the present invention under non-ideal imaging conditions.
[0042] like Figure 3 As shown, the process begins with the airborne image sensor acquiring the raw image. The image is then fed into the gradient analysis module (1.0), which generates the gradient magnitude statistical distribution, spatial orientation map, and dominant gradient direction and pixel clustering. The gradient magnitude statistical distribution is then fed into the image validity diagnosis module (2.0), which references the sharpness threshold database and outputs the image validity signal. The spatial orientation map is then fed into the feature credibility verification module (3.0), which outputs the feature credibility weights. Finally, the target decision module (4.0), upon receiving the image validity signal, feature credibility weights, and dominant gradient direction and pixel clustering, makes a target / background judgment based on its internal gradient-target mapping table and decision threshold database, and records the result in the inspection log.
[0043] Example 4: Before applying the system of this invention to the inspection task of detecting linear microcracks on solar cell array panels, an offline parameter and decision rule calibration procedure needs to be executed. This procedure is performed on a ground workstation that has pre-stored a calibration image library containing thousands of images. The images in the library cover solar cell panel samples with typical linear microcracks taken under different lighting and angle conditions, as well as pure background samples consisting of the sky or intact panel surfaces. After the calibration procedure is started, the gradient-target mapping table is first constructed. The system operator selects a subset of images containing the target from the calibration image library and manually outlines the images that clearly contain target features, specifically... The system first selects image blocks of linear microcracks. Then, it automatically runs the gradient analysis module on all selected image blocks and a large number of pure background image blocks in the image library, and records the dominant gradient direction and pixel clustering degree of each image block, thus forming a gradient feature database associated with the target type. Based on this database, the system determines the decision rules for each type of target through statistical analysis. By analyzing the gradient features of all image blocks marked as linear microcracks, it is found that the dominant gradient direction is concentrated in the direction perpendicular to the crack direction, and the pixel clustering degree is distributed in the range of 0.85 to 1.0. The system then stores the combination of this direction type and clustering degree range as a decision rule in the gradient-target mapping table.
[0044] After the gradient-target mapping table is generated, the system further calibrates the decision threshold used by the target decision module. This process utilizes the gradient feature database generated in the previous step, which already contains pixel clustering distribution data of target image blocks and background image blocks. The system multiplies these pixel clustering data with an initially set feature confidence weight, which is set to 1.0 here, to obtain two sets of weighted clustering statistical distributions. The technical consideration in setting the decision threshold is to balance the recall and precision of the detection. Its setting value is determined by analyzing the receiver operating characteristic curves of the two sets of weighted clustering distributions. In a specific case where the task requirement is to reduce false negatives, the value corresponding to the coordinate point on the curve that corresponds to a higher true positive rate can be selected as the decision threshold. This threshold and the constructed gradient-target mapping table are then fixed in the non-volatile memory of the UAV platform. The execution of this calibration procedure adapts the decision model and threshold parameters within the system to a specific detection task, providing a validated decision model and parameters for deploying the system of this invention in this specific application scenario.
[0045] Example 5: When the system of the present invention is deployed to a new work site with different background characteristics from the initial calibration environment, it can execute a field baseline adaptive adjustment procedure. In a specific application, a UAV that has completed offline calibration is assigned to perform structural inspection on a severely weathered concrete dam with a large number of irregular textures on its surface. The texture characteristics of the dam surface are different from the background sample used for offline calibration, which may affect the applicability of the original decision threshold.
[0046] Before the inspection mission officially began, the drone was set to environmental baseline acquisition mode. In this mode, the drone flew briefly over a section of the dam surface confirmed to be free of structural defects, following a planned flight path. During the flight, the onboard image sensor continuously acquired images, and the gradient analysis module and feature reliability verification module within the system operated, continuously calculating the weighted clustering degree values of the acquired background image patches. The target decision module then statistically analyzed these values until a preset number of sample data points were collected. After the acquisition was completed, the system performed statistical distribution analysis on these weighted clustering degree samples representing the background characteristics of the current working environment, calculating their mean μ. b With standard deviation σ b Based on the statistical analysis results, the system adaptively fine-tunes the decision threshold in the target decision module. The adjustment rule is to adjust the new decision threshold T... new Set as T new =μ b +k·σ b The coefficient k is a preset confidence factor, the value of which can be set according to the requirements of the task for the false alarm rate and the false alarm rate. Here, k is set to 3. This adjustment process makes the decision threshold match the background noise level measured on site, so that the system can adapt to the new background features without changing the embedded recognition rules in the gradient-target mapping table. After completing this baseline self-adaptive adjustment, the system switches to the standard operation mode and begins to perform the inspection task of the entire dam.
[0047] Example 6: When integrating the system of the present invention into a UAV platform equipped with a new type of airborne image sensor and using it to perform an inspection task requiring the identification of tiny targets, a system-level parameter tuning procedure needs to be executed before task deployment. This procedure aims to match the image block size and feature confidence verification logic within the system with the resolution of the new image sensor and the pixel size of the target features. The procedure first determines the image block size N used by the gradient analysis module. The technical consideration for setting this size is to balance the spatial resolution capability of the detection algorithm for tiny targets with the overall computational load of the system. The tuning process includes using the newly integrated UAV... The system captures a set of standard-resolution test images containing lines of varying pixel widths, corresponding to the minimum feature size of the target image. The system operator then sets a set of candidate values for N, including 8, 16, and 32, and processes the test images using the system. The system records the pixel clustering of the smallest recognizable line and the processing time per frame for each N value. Finally, the N value that minimizes the processing time per frame while effectively recognizing the target image is selected as the fixed parameter for this hardware configuration. After determining the image patch size N, the procedure then sets the confidence entropy threshold H in the feature confidence verification module. trust With noise entropy threshold H noise The calibration process uses two types of image datasets: one containing structures with strong spatial continuity, and the other containing stochastic gradient fields generated by high-frequency optical noise. The system calculates the continuity entropy H for all image patches in both datasets, forming two independent sets of entropy value statistical distributions. By analyzing these two distributions, H is... trust Set the upper limit of the distribution to cover 95% of the entropy values of the structured dataset, and simultaneously set H noise The two thresholds, set to cover 95% of the entropy values of the stochastic gradient field dataset, provide the feature credibility verification module with a judgment range to distinguish between real structures and random noise.
[0048] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A UAV-borne visual online detection and target recognition system, characterized in that, The system includes: A gradient analysis module is configured to divide an image acquired by an airborne image sensor into multiple image blocks, and for each image block, generate a gradient magnitude statistical distribution and a spatial orientation map composed of the dominant gradient directions of each image block based on pixel gradient calculation. An image validity diagnosis module, whose input is connected to a gradient analysis module, is configured to compare the proportion of high-amplitude gradient pixels in the gradient magnitude statistical distribution with a sharpness threshold stored in an onboard non-volatile memory, and generate an image validity signal characterizing whether the current image contains valid information based on the comparison result. A feature confidence verification module, whose input is connected to the gradient analysis module, is configured to generate a feature confidence weight that characterizes the authenticity of gradient information based on the analysis of the spatial relationship continuity between the dominant gradient directions of adjacent image blocks in the spatial orientation map. A target decision module, whose input is connected to a gradient analysis module, an image validity diagnosis module, and a feature credibility verification module, is configured to only combine the pixel clustering degree and feature credibility weight of each image patch with a gradient-target mapping table that is also pre-stored in an onboard non-volatile memory when an image validity signal is received indicating that the image contains valid information, and then match and determine the target type of the image patch from a gradient-target mapping table that is also pre-stored in an onboard non-volatile memory.
2. The UAV-borne visual online detection and target recognition system according to claim 1, characterized in that, The gradient analysis module is further configured to determine the dominant gradient direction and pixel clustering by performing voting statistics on the pixel gradient directions in four discrete directions within the image block. The pixel clustering is the ratio of the number of votes in the dominant gradient direction to the total number of votes.
3. The UAV-borne visual online detection and target recognition system according to claim 1, characterized in that, The image validity diagnosis module is further configured to generate an image validity signal indicating that the image does not contain valid information when the proportion of high-amplitude gradient pixels is determined to be lower than the sharpness threshold, and the target decision module will stop running after receiving the signal.
4. The UAV-borne visual online detection and target recognition system according to claim 1, characterized in that, The feature confidence verification module is further configured to: calculate the continuity entropy H of the spatial orientation map based on a set of transition probabilities p(i) between adjacent image blocks of the dominant gradient direction, where p(i) is the probability of transitioning from one dominant gradient direction to another; and generate feature confidence weights based on the calculation result of the continuity entropy H, where the value of the continuity entropy H is negatively correlated with the value of the feature confidence weights.
5. The UAV-borne visual online detection and target recognition system according to claim 1, characterized in that, The target decision module is further configured to: multiply the pixel clustering degree with the feature confidence weight to obtain a weighted clustering degree, and compare the weighted clustering degree with a decision threshold stored in an onboard non-volatile memory, using the comparison result as the basis for matching from the gradient-target mapping table.
6. The UAV-borne visual online detection and target recognition system according to claim 1, characterized in that, The gradient-target mapping table contains multiple entries, each defining a unique target type and a set of corresponding decision rules. The decision rules include the type of the dominant gradient direction and a weighted clustering range.
7. The UAV-borne visual online detection and target recognition system according to claim 2, characterized in that, The gradient analysis module uses the Sobel operator to calculate pixel gradients; and the four discrete directions include the horizontal and vertical directions and two diagonal directions.
8. The UAV-borne visual online detection and target recognition system according to claim 5, characterized in that, The target decision module is further configured to: when the weighted clustering degree is not higher than the decision threshold, the image patch is identified as a background region, and no subsequent target type matching and output are performed, so as to reduce the overall computational load of the system.
9. The UAV-borne visual online detection and target recognition system according to claim 3, characterized in that, The percentage of high-amplitude gradient pixels is obtained by counting the number of pixels with gradient amplitudes greater than the sharpness threshold and then dividing by the total number of pixels.
10. The UAV-borne visual online detection and target recognition system according to claim 1, characterized in that, Physically, the system is an embedded computing unit integrated into the UAV platform. The logical functions of the gradient analysis module, image validity diagnosis module, feature credibility verification module, and target decision module are all executed by one or more microcontroller units configured with firmware. The microcontroller units are connected to the airborne image sensor and airborne non-volatile memory.
Citation Information
Cited By
Crack visual identification method for abnormal stress of bidirectional eccentric compression member
CN121962861A