A fusion calibration method for UAV target detection based on visual and radio features
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]针对现有技术存在的不足,本发明的目的在于提供基于视觉及无线电特征的无人机目标检测融合校准方法,能够解决现有技术中无人机多源融合定位在复杂低空作业场景下,融合层次局限于坐标级而无法区分不同视觉目标特征可信度差异、未充分利用无线电定位数值准确但向量性差的差异化特性、以及缺乏独立外部基准对特征权重进行动态校准的问题,本发明通过深入特征层面的一致性匹配与逐特征赋权,并引入外部安防相机作为独立基准进行校准,实现了无人机在复杂低空环境下自身定位精度与鲁棒性的显著提升
[0015]本发明的有益效果:将融合层次从传统的坐标级下沉至特征级,通过对视觉图像中不同目标特征的跨帧变化状态分别与无线电运动趋势进行一致性匹配,为每个目标特征独立分配权重系数,有效区分了可靠特征与不可靠特征对定位的贡献,避免了不可靠特征误差的无差别引入,再利用无线电定位坐标数值准确的优势进行特征一致性校验,同时规避其位移向量噪声大的劣势,仅在特征匹配阶段发挥其校验标尺作用,不参与最终定位坐标融合;另外,引入外部安防相机的基准轨迹点作为独立参照系,通过虚拟锚点行为特征因子识别无人机运动场景,并依据场景和偏差向量对各目标特征的权重系数进行动态校准,形成了偏差度量、场景感知以及权重调整的完整路线,使无人机在脱离外部视觉覆盖后仍能保持高精度定位能力。
Smart Images

Figure CN122384838B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of UAV positioning and navigation and multi-source information fusion technology, and more specifically to a UAV target detection fusion calibration method based on visual and radio features. Background Technology
[0002] With the widespread application of drones in power line inspection, construction site monitoring, and emergency rescue, increasingly higher demands are being placed on the precise positioning capabilities of drones in complex low-altitude environments. However, in actual operations, drones face multiple positioning challenges. For example, in scenarios such as power transmission line corridors, mountain valleys, and urban building complexes, signals are easily blocked by towers, mountains, and buildings, or multipath reflections occur, leading to a sharp decline in positioning accuracy or even complete failure. Airborne inertial measurement units accumulate drift errors during long-term operation, and these errors increase over time, making it impossible to maintain accurate positioning capabilities independently. The reliability of visual feature extraction and tracking by airborne visual sensors is significantly reduced under conditions such as drastic changes in lighting, missing textures, rapid movement causing image blurring, or the presence of a large number of repetitive textures in the scene. Airborne radio positioning signals are easily affected by multipath effects, co-channel interference, and changes in antenna directivity in complex electromagnetic environments, resulting in insufficient stability and directional accuracy of positioning results.
[0003] To address the limitations of single sensors, multi-source information fusion positioning technology is currently widely used to solve the aforementioned problems. However, in practical applications, existing multi-source fusion positioning methods cannot distinguish the reliability differences in the contribution of different target features in images to positioning in complex low-altitude operation scenarios. This leads to the indiscriminate introduction of errors from unreliable features into the fusion result. Furthermore, by treating radio positioning as a common coordinate source for equal fusion, the problem of large displacement noise cannot be avoided. In addition, when the weight coefficients of visual target features become unsuitable due to environmental changes, the UAV itself cannot perceive and correct the deviation. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a UAV target detection fusion calibration method based on visual and radio features. This method solves the problems in existing technologies where multi-source fusion positioning of UAVs in complex low-altitude operating scenarios is limited to the coordinate level and cannot distinguish the differences in the reliability of different visual target features, fails to fully utilize the differential characteristics of radio positioning, which is accurate but has poor vector properties, and lacks an independent external benchmark for dynamic calibration of feature weights. The present invention achieves a significant improvement in the self-positioning accuracy and robustness of UAVs in complex low-altitude environments by deepening the consistency matching and feature-by-feature weighting at the feature level and introducing an external security camera as an independent benchmark for calibration.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for fusion calibration of UAV target detection based on visual and radio features includes the following steps: The target feature extraction step involves acquiring images of the surrounding environment using a visual camera mounted on a drone, identifying target features in the images, and extracting visual target features of multiple preset categories. The visual target features include at least static and dynamic ground features. The feature matching and weighting step tracks the cross-frame change state of each target feature in adjacent frame images, and performs consistency matching between the change state of each target feature and the UAV motion trend indicated by the radio positioning coordinates to obtain the corresponding consistency score. Then, the weight coefficient of each target feature is determined based on the consistency score. The weight coefficient corresponding to each target feature is used as the fusion weight to perform weighted fusion of its cross-frame change state and accumulate frame by frame to obtain the UAV visual positioning coordinates. The cross-source deviation calculation step involves aligning the reference trajectory points of the UAV located by the external camera with the visual positioning coordinates and then calculating the deviation vector. The weighted dynamic calibration step identifies the current motion scene of the UAV based on the virtual anchor point behavior feature factors generated from the reference trajectory points, and dynamically adjusts the weight coefficients of each target feature according to the motion scene and the deviation vector to obtain the calibrated weight coefficients. The calibration positioning output step involves re-executing the feature matching and weighting step with the calibrated weight coefficients and outputting the calibrated UAV visual positioning coordinates.
[0006] Furthermore, the feature matching and weighting step also includes a cross-frame feature matching strategy. The cross-frame feature matching strategy includes associating and tracking the same target feature in adjacent frame images, obtaining the pixel displacement vector and scale change rate of the target feature as the cross-frame change state; then calculating the UAV motion trend vector based on the radio positioning coordinate difference between adjacent time moments, comparing the motion direction indicated by the pixel displacement vector of each target feature with the UAV motion trend vector for directional consistency, and determining the consistency score based on the directional deviation.
[0007] Furthermore, the feature matching and weighting step also includes a weighted cumulative localization strategy. The weighted cumulative localization strategy includes: for each target feature, calculating the inter-frame motion estimate of the UAV implied by the feature based on its pixel displacement vector and scale change rate; using the weight coefficients corresponding to each target feature as fusion weights, performing a weighted average of the inter-frame motion estimates of all target features to obtain the fused motion estimate of the current frame; and accumulating the fused motion estimate onto the visual localization coordinates of the previous frame to obtain the UAV visual localization coordinates of the current frame.
[0008] Furthermore, the virtual anchor point behavior feature factors include angular velocity amplitude, directional oscillation rate, angular monotonicity, area change rate, and area linear fitting slope. The weight dynamic calibration step includes a feature factor calculation strategy, which includes calculating the angular velocity amplitude based on the rate of change of the angle between adjacent trajectory points in the baseline trajectory point sequence; calculating the directional oscillation rate based on the frequency of trajectory turning direction changes within a statistical sliding window; calculating the angular monotonicity based on whether the turning angle sequence within the evaluation sliding window continuously increases or decreases; calculating the area change rate based on the change in the area of the polygon enclosed by the trajectory point sequence and the bottom edge of the image per unit time; and calculating the area linear fitting slope based on the slope obtained by linear fitting the polygon area sequence at each time point within the sliding window.
[0009] Furthermore, the dynamic weight calibration step also includes a scenario-based weight adjustment strategy, which includes identifying the current motion scenario of the UAV based on virtual anchor point behavior feature factors. The motion scenario includes at least a high-dynamic turning scenario, a visually degraded approach scenario, and a stable hovering scenario. When a high-dynamic turning scene is identified, the reliability of the radio positioning coordinate vector is determined to be reduced, and the feature weights positively correlated with the radio positioning coordinates are decreased. When a visually degraded approach scene is identified, the tracking quality of visual target features is determined to be reduced, and the weights of dynamic ground features are decreased while the weights of static ground features are increased. When a stable hovering scene is identified, if the magnitude of the deviation vector exceeds a preset threshold, a systematic cumulative error is determined, and the weight coefficients of each target feature are iteratively optimized to minimize the magnitude of the deviation vector.
[0010] Furthermore, the scenario-specific weight adjustment strategy also includes identifying a high-dynamic turning scenario when the angular velocity amplitude exceeds a preset angular velocity threshold and the angle monotonicity is lower than a preset monotonicity threshold; identifying a visual degradation approach scenario when the absolute value of the area linear fitting slope exceeds a preset slope threshold and the slope is positive; and identifying a stable hovering scenario when the angular velocity amplitude is lower than a preset angular velocity threshold, the direction oscillation rate is lower than a preset oscillation rate threshold, and the area change rate is lower than a preset area change rate threshold.
[0011] Furthermore, it also includes an environment-adaptive filtering step, which includes quantifying the contour change amplitude and deformation degree of the same target feature in adjacent frame images, and then using the contour change amplitude and deformation degree as the judgment criteria to determine whether the target feature is affected by environmental factors, marking or removing target features that are determined to be affected by environmental factors, and restricting the consistency matching and weight allocation of the target feature.
[0012] Furthermore, the environment adaptive screening step also includes a feature tracking strategy, which includes statistically analyzing the tracking continuity and matching confidence of each target feature across multiple consecutive frames. The tracking continuity reflects the trajectory continuity of the target feature between adjacent frames, and the matching confidence reflects the feature descriptor similarity of the target feature between adjacent frames. Based on the tracking continuity and matching confidence, it is determined whether to pause feature matching weighting of the target feature in the current frame.
[0013] Furthermore, when the trajectory of a target feature is interrupted between adjacent frames or the similarity of the feature descriptors is lower than a preset similarity threshold, the feature matching weighting of the target feature in the current frame is suspended. When the trajectory of the suspended target feature is restored to continuity between subsequent adjacent frames or the similarity of the feature descriptors reaches the preset similarity threshold, the feature matching weighting of the target feature in the current frame is restored.
[0014] Furthermore, the static ground features include one or more of power poles, utility poles, buildings, and insulators, while the dynamic ground features include one or more of trees, clouds, pedestrians, and moving vehicles.
[0015] The beneficial effects of this invention are as follows: The fusion level is lowered from the traditional coordinate level to the feature level. By consistently matching the cross-frame changes of different target features in the visual image with the radio motion trend, and independently assigning weight coefficients to each target feature, the contribution of reliable and unreliable features to positioning is effectively distinguished, avoiding the indiscriminate introduction of unreliable feature errors. Furthermore, the accuracy of radio positioning coordinate values is utilized for feature consistency verification, while avoiding its disadvantage of high displacement vector noise. It only plays a verification role in the feature matching stage and does not participate in the final positioning coordinate fusion. In addition, the reference trajectory points of external security cameras are introduced as an independent reference system. Virtual anchor point behavioral feature factors are used to identify the UAV's motion scene, and the weight coefficients of each target feature are dynamically calibrated based on the scene and deviation vector. This forms a complete route for deviation measurement, scene perception, and weight adjustment, enabling the UAV to maintain high-precision positioning capabilities even after leaving external visual coverage. Attached Figure Description
[0016] Figure 1 This is the overall flowchart of the present invention; Figure 2 This is a flowchart of the environmental adaptive screening step in this invention; Figure 3 This is a schematic diagram of the external camera positioning of the drone in this invention. Detailed Implementation
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.
[0018] This invention discloses a UAV target detection fusion calibration method based on visual and radio features, which solves the problems of shallow fusion layer, inability to distinguish feature reliability, and lack of effective dynamic calibration mechanism in existing technologies. Specifically, firstly, instead of directly fusing visual positioning results with radio positioning results at the coordinate level, it delves into the visual feature level, using the motion trend provided by radio positioning as a "ruler" to perform consistency evaluation on each trackable target feature in the image and assign differentiated weights to features with different reliability. Secondly, it introduces reference trajectory points provided by external security cameras or other third-party vision systems as absolute references. By calculating the deviation between the visual positioning coordinates and this reference, and combining it with behavioral feature factors derived from the reference trajectory points, it identifies the UAV's motion scene and dynamically calibrates the feature weight allocation strategy accordingly. Through the above methods, a complete route from feature extraction and weight assignment to dynamic calibration is achieved, significantly improving the positioning accuracy and robustness of UAVs in complex low-altitude environments.
[0019] Example 1 This embodiment describes a scenario in a typical power transmission line inspection task where a drone flies along a preset route and needs to accurately detect and locate targets such as power poles and insulators. In this scenario, the drone may face situations such as GPS signals being weakened due to pole blockage and the coexistence of poles (static ground features) and moving vehicles (dynamic ground features) below in the visual image.
[0020] like Figure 1 As shown, the fusion calibration method of the present invention includes a target feature extraction step, a feature matching and weighting step, a cross-source bias calculation step, a weight dynamic calibration step, and a calibration positioning output step. In the target feature extraction step, the forward-looking or downward-looking vision camera on the UAV collects image sequences of the surrounding environment in real time at a set frame rate. After acquiring the current frame image, preprocessing including denoising, distortion correction, and image enhancement is performed to improve the quality of subsequent feature extraction. Then, the target features are identified by a target detection network based on deep learning, which is specifically trained for power inspection scenarios. The network has a preset detection category library containing at least two types of visual target features: static ground features and dynamic ground features. In this embodiment, static ground features may include the tower body, crossarm, insulator string, power pole, transmission line, and substation buildings, etc. Static ground features can be regarded as an absolutely stationary reference system in a short period of time. Dynamic ground features may include tree branches and leaves swaying in the wind, vehicles driving on the road below the inspection path, the shadow edges of clouds drifting in the sky, and even moving boats in the water.
[0021] This network identifies and outputs all detected target features for each frame of the image. For each detected target feature, it records its category, confidence level, geometric center position, feature contour, and feature descriptor for subsequent inter-frame matching. The feature descriptor is such as the SIFT feature vector. In actual shooting, the tower body (static), insulator string (static), and moving truck (dynamic) can be identified simultaneously in the frame-shifting image. Although the image is a static image, it can be preliminarily determined whether it is static or dynamic based on the target feature category.
[0022] The feature matching and weighting step, after successfully extracting multiple target features, uses the motion trend provided by radio positioning coordinates to independently evaluate and weight the credibility of each visual feature. Specifically, when implementing the cross-frame feature matching strategy, optical flow or feature point matching algorithms are used to track each target feature between the current frame and the previous frame. For each successfully tracked feature, its cross-frame change state is calculated. This state mainly includes two quantitative indicators: pixel displacement vector and scale change rate. The pixel displacement vector reflects the direction and distance of movement of the center point of the target feature from the previous frame to the current frame in the pixel coordinate system. Let the coordinates of the feature center in the previous frame be... The center coordinates of the current frame are Then the pixel displacement vector is For example, when a drone flies forward, the tower features in the image may expand and move outward from the image center. The scale change rate reflects the proportion of change in the bounding box area of the target feature between frames, i.e., the scale change rate. ,in These are the pixel areas of the feature bounding boxes in the previous and current frames, respectively. When the drone approaches a target, the scale of the target increases, and when it moves away, it decreases.
[0023] Simultaneously, the positioning coordinates output by the UAV's onboard radio positioning equipment are acquired, and a three-dimensional UAV motion trend vector is obtained by calculating the difference between radio positioning coordinates at adjacent times. This vector indicates the actual direction and approximate distance of the UAV's movement in the world coordinate system; then the pixel displacement vector of the target feature is... Direction vector from pixel coordinate system to camera coordinate system Then, compare its direction consistency with the drone's motion trend vector and calculate the angle between the two vectors. Consistency score ,in For a preset tolerance threshold for directional deviation, the smaller the angle, the more consistent the motion direction implied by the visual feature is with the actual motion direction measured by radio. The consistency score is... The higher the score, the closer it is to 1. Conversely, if the pixel displacement direction of a dynamic object (a vehicle moving in the opposite direction or a person walking in the opposite direction) is significantly different from the drone's movement direction, its consistency score will be very low. In this way, radio positioning becomes a "benchmark" for verifying the reliability of features.
[0024] When implementing the weighted cumulative localization strategy, the weight coefficients are determined by the consistency score. Specifically, for each target feature... Its weighting coefficient By its consistency score Softmax normalization is performed to obtain According to its pixel displacement vector and scale transformation rate Combined with the camera's intrinsic parameters (focal length) Principal point coordinates ) and the drone's flight altitude at the time Alternatively, depth information can be used to infer the inter-frame motion estimation of the UAV implied by this feature through geometric relationships. That is, the rate of change on one scale. The displacement in the depth direction can be deduced. Pixel displacement can be used to deduce displacement on the horizontal plane. .
[0025] The reliability of inter-frame motion estimates provided by different features varies. Therefore, the consistency score of each feature, after normalization, is used as its weight coefficient. The higher the weight coefficient, the more reliable the motion estimate provided by that feature in the current environment. Subsequently, the inter-frame motion estimates of all features are... With its corresponding weight coefficient A weighted average is then performed to obtain a robust fusion motion estimate. Finally, the calculated fused motion estimates are accumulated into the UAV visual positioning coordinates solved in the previous frame. This allows us to obtain the UAV's visual positioning coordinates for the current frame. This process constructs a visual odometry based on feature-weighted voting. Its beneficial effect is that when there are a large number of inconsistent dynamic interference features in the image, their weights are automatically reduced, thereby suppressing error accumulation and ensuring the smoothness and accuracy of visual positioning. For example, in flight, highly consistent static tower features are given high weights and dominate the final positioning result; while low-consistency moving vehicle features have extremely low weights, and their interference with positioning is effectively filtered out.
[0026] In the cross-source deviation calculation step, to correct for the cumulative drift that may occur during the visual positioning process, this invention introduces an independent external security camera system. This security camera can be deployed on the top of a pole or other building at a high point in the inspection area. Its world coordinates are known through calibration, and it can use image recognition algorithms to locate and track the pixel position of the drone in the image in real time. Figure 3 As shown, the reference trajectory points of the UAV in the real coordinate system are then calculated through coordinate transformation. .
[0027] The system receives the reference trajectory points from an external camera via a communication link. Since the visual positioning coordinates and the reference trajectory points may be located in different world coordinate systems (i.e., the visual positioning origin is at the starting point, while the security camera coordinates are absolute geographic coordinates), it is first necessary to align the coordinates using the coordinate sequence within the most recent time window and calculate the optimal transformation parameters and rotation matrix between the two coordinate systems. Translation vector After completing the coordinate alignment, calculate the aligned UAV visual positioning coordinates. With reference trajectory points provided by security cameras Deviation vector between This deviation vector includes not only the cumulative error magnitude of visual positioning It also includes information about the direction of the error, such as the deviation vector possibly being offset by a certain distance in the northeast direction.
[0028] The dynamic weight calibration step is triggered once a stable deviation vector is obtained. Specifically, when implementing the feature factor calculation strategy, the original baseline trajectory point coordinates are not used directly. Instead, the sequence is subjected to secondary mining to generate a series of virtual anchor point behavioral feature factors. These factors can highly abstractly describe the current motion mode of the UAV, including angular velocity amplitude. directional oscillation rate Angular monotonicity Area change rate and the slope of the area linear fitting angular velocity amplitude It is calculated based on the rate of change of the angle formed by the line connecting three adjacent points in the reference trajectory point sequence. Let the three adjacent trajectory points be... Calculate vector (representing the displacement from the previous moment to the current moment) and (representing the displacement from the current moment to the next moment), and its included angle. This refers to the steering angle and the amplitude of the angular velocity. , The time interval reflects the speed at which the drone turns; directional oscillation rate It involves statistically analyzing the frequency of positive and negative changes in the drone's flight direction (trajectory turning angle) within a given time window, and recording each turning angle. The directional oscillation rate of forward or reverse rotation. , A higher value indicates more frequent left-right swaying of the flight direction, reflecting the degree of flight stability. High-frequency oscillations usually mean that the drone is struggling against wind disturbances or performing fine hovering; angle monotonicity It is used to distinguish between stable hovering and random jitter, based on the sequence of turning angles within a sliding window. Perform least-squares linear fitting to obtain the slope of the fitted line. If the absolute value of the slope The coefficient of determination for linear fit that is greater than the preset monotonic threshold If the value is close to 1, it is considered to have high monotonicity; otherwise, it is considered to have low monotonicity. (Area change rate) It is used to quantify the radial velocity of a drone relative to an external camera, by selecting a fixed spatial reference point. This can be the position of the security camera in the world coordinate system or any chosen origin. At time, the reference trajectory point sequence is set from the starting time. up to the current moment With reference point Connect the elements sequentially to form a dynamic polygon, and calculate the total area of the polygon. In practical discrete sampling systems, the area of a polygon can be determined by comparing adjacent trajectory points with a reference point. The area of the multiple triangles formed is summed to obtain the result, followed by the rate of change of area. Calculated using discrete difference method: This value reflects the rate of change of the polygon's area over time; a positive value... This means the drone is relative to the reference point As it moves away from or nears; the slope of the linear fit of the area. To identify visually degraded proximity scenes, similar to angular monotonicity, a linear fit is performed on the area sequence calculated within a sliding window, and the resulting slope value is... ,if A significantly positive slope exceeding the preset slope threshold indicates that the drone is continuously and stably approaching the large structure, which is a typical characteristic of visually degraded approach scenes.
[0029] When implementing the scenario-based weight adjustment strategy, the above factors are compared with preset thresholds to automatically identify the current motion scenario of the drone and perform calibration based on the scenario and deviation vector. Including scenario one, a high-dynamic turning scenario, when the angular velocity amplitude is detected... Exceeding a higher threshold And the angle is monotonous Below a low threshold When the direction change is monotonic, it is determined that the UAV is entering a rapid turning maneuver. In this scenario, due to antenna directivity and multipath effects, the reliability of the displacement vector provided by the radio positioning device will decrease sharply. Therefore, the calibration strategy is to reduce the feature weights that are positively correlated with the radio positioning coordinates, i.e., by introducing a turning scenario coefficient. To recalculate the consistency score: ,in Directly with Negative correlation, the faster the turn The smaller, this will lead to In the calculation, the influence of consistency scores based on radio scales is significantly reduced, and greater reliance is placed on the continuity and smoothness of visual feature tracking itself, thereby avoiding the introduction of unreliable radio noise into the localization results.
[0030] Scenario 2: Visual degradation is similar to the scene where the slope of the linear fitting of the area is recognized. The absolute value exceeds the threshold When the slope is positive, it indicates that the drone is continuously approaching a large object, such as a power transmission tower. At this time, the visual image will be rapidly magnified, causing the texture details of the target to become blurred and feature points to quickly move out of the field of view, resulting in a severe decrease in tracking quality. The calibration strategy is to reduce the weight of dynamic features and increase the weight of static features. Specifically, it suppresses the contribution of dynamic or volatile features such as leaves and cloud shadows, while strengthening the dependence on rigid static features such as building edges and tower structures. That is, it introduces a category weight factor based on its classification for each feature. For static ground features such as insulators, Set an enhancement factor greater than 1 for dynamic ground features such as leaves. Set to a suppression coefficient less than 1, the final fusion weights are adjusted to... These high-contrast static features are used to maintain the stability of positioning and prevent positioning jumps caused by rapid changes in texture.
[0031] Scenario 3, stable hovering scenario, when the angular velocity amplitude is detected. directional oscillation rate and area change rate When all values are below their respective low thresholds, the drone is determined to be in a stable hovering or near-hovering state. At this point, if the magnitude of the deviation vector obtained from the cross-source deviation calculation step is... Exceeding the preset tolerance threshold If the visual positioning system exhibits a slow, systematic accumulation error, the calibration strategy involves iteratively optimizing the weight coefficients to minimize the magnitude of the deviation vector. Specifically, a gradient descent optimizer is initiated, with the objective function set as follows: The optimization variables are the weighted coefficients of various features or the weight ratio of consistency scores. The weight coefficients of each type of feature are adjusted by small steps, such as slightly increasing the weight of static features with absolute scale constraints, and the change of the deviation vector at the next moment is observed. If the magnitude decreases, the adjustment continues in the same direction; if it increases, the adjustment is reversed. This method is an optimization process, the goal of which is to gradually pull the visual positioning coordinates back to the reference trajectory provided by the security camera, thereby eliminating the accumulated error online.
[0032] Among them, the calibration positioning output step and the calibrated weight coefficients output by the weight dynamic calibration step are included. These optimized strategy parameters are then reapplied to the feature matching weighting step in the next frame image processing. Specifically, in the weighted cumulative localization strategy of subsequent frames, these updated weight coefficients will be used to calculate the fused motion estimate, thereby outputting the calibrated UAV visual localization coordinates with higher accuracy.
[0033] Example 2 This embodiment is based on embodiment 1, such as... Figure 2As shown, it also includes an environment-adaptive screening step. This step, located between the target feature extraction step and the feature matching and weighting step, is used to address scenarios where radio signals are subject to severe electromagnetic interference and visual images are simultaneously affected by environmental disturbances, such as gimbal shaking caused by wind, and swaying of edge modules and leaves caused by fog. In the scenario of power line inspection, UAVs face two levels of problems. First, the coordinates output by the airborne radio positioning equipment contain a lot of noise, and the signal-to-noise ratio of its motion trend vector is extremely low, making it unsuitable as a "ruler" for consistent comparison of all features. Second, environmental disturbances can cause non-rigid deformation and shaking of static ground features in the image (such as insulator strings and tower crossarms) in the image sequence. If these static features, which should be highly reliable, are incorrectly assigned low weights due to environmental disturbances, it will seriously impair positioning accuracy. Therefore, the environment-adaptive screening step is set up to identify and screen out high-reliability, less affected features by environmental disturbances by analyzing the stability of the visual features themselves, without relying on the radio ruler.
[0034] Specifically, this includes a quantitative evaluation and screening sub-strategy. For each visual feature identified in the target feature extraction step, especially static ground features, the stability of its contour is quantitatively analyzed in the tracked consecutive frame images. That is, for each target feature, its contour Hu moments are extracted as shape descriptors. Hu moments are a set of moment feature quantities constructed from normalized central moments that have translation, rotation, and scale invariance, and can effectively describe the overall shape of a region. The current frame is denoted as... Chinese characteristics The Hu moment vector is .
[0035] Next, the Hu moment difference of the uniform feature contours between adjacent frames is calculated to quantify the degree of contour deformation and the difference. Defined as the Mahalanobis distance or Euclidean distance between two Hu moment vectors: For rigid objects like towers under ideal conditions, their outlines should remain almost unchanged between adjacent frames when there is no disturbance. However, when high-frequency jitter of the pan / tilt unit, fog causing edge blurring, or strong electromagnetic interference coupling into the imaging system and generating image noise occur, even the extracted outlines of static rigid objects will undergo drastic and irregular deformation, leading to… Significantly increased.
[0036] To avoid the impact of single-frame noise, calculations are performed within a sliding window. The cumulative instability index within , It can be the mean or sum of the deformation measures between frames within the window. This index comprehensively reflects the overall stability of the feature over a period of time. By introducing Hu moments and inter-frame difference measures, this purely visual and quantitative contour stability assessment method, which does not rely on any external benchmarks, can effectively distinguish between rigid body features that maintain shape stability under environmental disturbances and unstable features that undergo drastic shape changes.
[0037] Based on the contour stability assessment, the continuity of feature tracking and matching confidence are further integrated to construct a comprehensive tracking quality score. First, calculate the proportion of times this feature was successfully tracked in the most recent historical frames, and let the window length be... The number of frames successfully tracked within the range is Then track continuity score The criteria for successful tracking are: firstly, feature matching must uniquely and explicitly match the feature near the predicted location; secondly, for each successful feature match, the ratio of the minimum Hamming distance to the second minimum Hamming distance between its feature descriptors is recorded. The smaller this ratio, the stronger the uniqueness of the surface match and the higher the matching confidence. The mean of the matching confidence within the window is calculated and normalized to... Interval as matching confidence score Finally, contour stability, tracking continuity, and matching confidence are weighted and fused to obtain a comprehensive tracking quality score for each feature. By comprehensively tracking the quality score, features lost due to brief occlusion or mismatched due to motion blur will have their scores reasonably reduced.
[0038] Based on the calculated comprehensive tracking quality score Each feature is filtered and pre-weighted, and a quality threshold is set. For rating The features identified are labeled as high-reliability features. These features typically correspond to rigid features in the image that have clear textures, stable outlines, and are minimally affected by environmental disturbances, such as clearly defined distant towers or sharp edges of buildings. These features are crucial for scoring. The characteristics are labeled as environmentally affected features, which may include insulators with blurred edges due to fogging, tree branches with distorted outlines due to wind-induced shaking, or textured areas with pixel noise due to electromagnetic interference.
[0039] For objects marked as having environmentally affected features, their initial weight coefficients are set to a minimum or zero before proceeding to the subsequent feature matching and weighting steps. This effectively "masks" low-quality features, preventing them from participating in subsequent radio consistency comparisons and weighted fusion localization. Conversely, high-reliability features retain their right to participate normally in subsequent calculations. When a low-quality feature's quality score is affected across multiple consecutive frames... When the weight rises above the threshold again, the suppression will be automatically lifted and the weight will be restored.
[0040] In scenarios where the radio scale fails, this step eliminates reliance on external signals. Instead, it filters out reliable visual anchors based on the inherent stability and tracking continuity of the visual features themselves. This is equivalent to performing a "pre-screening" based on the inherent consistency of vision before fusion localization, ensuring that only truly stable features can be included in the subsequent localization calculation. This effectively solves the problem that traditional methods treat all features lightly in complex and disturbed environments, leading to a large amount of noise being injected into the fusion model.
[0041] Furthermore, when the system detects an interruption in the trajectory of a target feature between adjacent frames—that is, when the predicted location cannot match a candidate feature, or when the similarity of its best-matching feature descriptor is lower than a preset similarity threshold—it does not immediately mark the feature as "missing" and remove it from the map. Instead, the system temporarily "pauses" the matching and weighting operation of that feature in the current frame. When calculating the weighted fusion motion estimation of the current frame, this feature will not provide any "votes." For paused features, they are not completely abandoned. Instead, their historical motion information before being paused (such as velocity and acceleration) is used to predict their expected position and appearance scale in the current frame through a Kalman filter. Feature search continues within the expected neighborhood. If a "predicted" version of the feature is successfully matched in a subsequent frame, and its confidence recovers to above the threshold, the system will automatically "restore" the feature, allowing it to re-participate in feature matching and weighting in subsequent frames.
[0042] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for fusion calibration of UAV target detection based on visual and radio features, characterized in that: Includes the following steps: The target feature extraction step involves acquiring images of the surrounding environment using a visual camera mounted on a drone, identifying target features in the images, and extracting visual target features of multiple preset categories. The visual target features include at least static and dynamic ground features. The feature matching and weighting step tracks the cross-frame change state of each target feature in adjacent frame images, and performs consistency matching between the change state of each target feature and the UAV motion trend indicated by the radio positioning coordinates to obtain the corresponding consistency score. Then, the weight coefficient of each target feature is determined based on the consistency score. The weight coefficient corresponding to each target feature is used as the fusion weight to perform weighted fusion of its cross-frame change state and accumulate frame by frame to obtain the UAV visual positioning coordinates. The cross-source deviation calculation step involves aligning the reference trajectory points of the UAV located by the external camera with the visual positioning coordinates and then calculating the deviation vector. The weighted dynamic calibration step identifies the current motion scene of the UAV based on the virtual anchor point behavior feature factors generated from the reference trajectory points, and dynamically adjusts the weight coefficients of each target feature according to the motion scene and the deviation vector to obtain the calibrated weight coefficients. The calibration positioning output step involves re-executing the feature matching and weighting step with the calibrated weight coefficients and outputting the calibrated UAV visual positioning coordinates.
2. The UAV target detection fusion calibration method based on visual and radio features according to claim 1, characterized in that: The feature matching and weighting step also includes a cross-frame feature matching strategy. The cross-frame feature matching strategy includes associating and tracking the same target feature in adjacent frame images, obtaining the pixel displacement vector and scale change rate of the target feature as the cross-frame change state; then calculating the UAV motion trend vector based on the radio positioning coordinate difference between adjacent time moments, comparing the motion direction indicated by the pixel displacement vector of each target feature with the UAV motion trend vector for directional consistency, and determining the consistency score based on the directional deviation.
3. The UAV target detection fusion calibration method based on visual and radio features according to claim 2, characterized in that: The feature matching and weighting step also includes a weighted cumulative localization strategy. The weighted cumulative localization strategy includes: for each target feature, calculating the inter-frame motion estimate of the UAV implied by the feature based on its pixel displacement vector and scale change rate; using the weight coefficients corresponding to each target feature as fusion weights, performing a weighted average of the inter-frame motion estimates of all target features to obtain the fused motion estimate of the current frame; and accumulating the fused motion estimate onto the visual localization coordinates of the previous frame to obtain the UAV visual localization coordinates of the current frame.
4. The UAV target detection fusion calibration method based on visual and radio features according to claim 1, characterized in that: The virtual anchor point behavior feature factors include angular velocity amplitude, directional oscillation rate, angular monotonicity, area change rate, and area linear fitting slope. The weight dynamic calibration step includes a feature factor calculation strategy, which includes: calculating the angular velocity amplitude based on the rate of change of the angle between adjacent trajectory points in the baseline trajectory point sequence; calculating the directional oscillation rate based on the frequency of trajectory turning direction changes within a statistical sliding window; calculating the angular monotonicity based on whether the turning angle sequence within the evaluation sliding window continuously increases or decreases; calculating the area change rate based on the change in the area of the polygon enclosed by the trajectory point sequence and the bottom edge of the image per unit time; and calculating the area linear fitting slope based on the slope obtained by linear fitting the polygon area sequence at each time point within the sliding window.
5. The UAV target detection fusion calibration method based on visual and radio features according to claim 4, characterized in that: The dynamic weight calibration step also includes a scenario-based weight adjustment strategy, which includes identifying the current motion scenario of the drone based on virtual anchor point behavioral feature factors. The motion scenario includes at least a high-dynamic turning scenario, a visually degraded approach scenario, and a stable hovering scenario. When identified as a high-dynamic turning scene, the reliability of the vector for determining radio positioning coordinates decreases, and the feature weights positively correlated with radio positioning coordinates are reduced. When a scene is identified as visually degraded, the tracking quality of visual target features is determined to be reduced, and the weight of dynamic features is reduced while the weight of static features is increased. When a stable hovering scenario is identified, if the magnitude of the deviation vector exceeds a preset threshold, it is determined that there is a systematic cumulative error, and the weight coefficients of each target feature are iteratively optimized to minimize the magnitude of the deviation vector.
6. The UAV target detection fusion calibration method based on visual and radio features according to claim 5, characterized in that: The scenario-specific weight adjustment strategy further includes identifying a high-dynamic turning scenario when the angular velocity amplitude exceeds a preset angular velocity threshold and the angle monotonicity is lower than a preset monotonicity threshold; identifying a visual degradation approach scenario when the absolute value of the area linear fitting slope exceeds a preset slope threshold and the slope is positive; and identifying a stable hovering scenario when the angular velocity amplitude is lower than a preset angular velocity threshold, the directional oscillation rate is lower than a preset oscillation rate threshold, and the area change rate is lower than a preset area change rate threshold.
7. The UAV target detection fusion calibration method based on visual and radio features according to any one of claims 1-6, characterized in that: It also includes an environment-adaptive filtering step, which includes quantifying the contour change amplitude and deformation degree of the same target feature in adjacent frame images, and then using the contour change amplitude and deformation degree as the judgment criteria to determine whether the target feature is affected by environmental factors, marking or removing the target feature that is determined to be affected by environmental factors, and restricting the consistency matching and weight allocation of the target feature.
8. The UAV target detection fusion calibration method based on visual and radio features according to claim 7, characterized in that: The environment adaptive screening step also includes a feature tracking strategy, which includes statistically analyzing the tracking continuity and matching confidence of each target feature in multiple consecutive frames. The tracking continuity reflects the trajectory continuity of the target feature between adjacent frames, and the matching confidence reflects the feature descriptor similarity of the target feature between adjacent frames. Based on the tracking continuity and matching confidence, it is determined whether to pause the feature matching weighting of the target feature in the current frame.
9. The UAV target detection fusion calibration method based on visual and radio features according to claim 8, characterized in that: When the trajectory of a target feature is interrupted between adjacent frames or the similarity of the feature descriptors is lower than a preset similarity threshold, the feature matching weighting of the target feature in the current frame is suspended. When the trajectory of the suspended target feature is restored to continuity between subsequent adjacent frames or the similarity of the feature descriptors reaches the preset similarity threshold, the feature matching weighting of the target feature in the current frame is restored.
10. The UAV target detection fusion calibration method based on visual and radio features according to claim 1, characterized in that: The static ground features include one or more of power poles, utility poles, buildings, and insulators, while the dynamic ground features include one or more of trees, clouds, pedestrians, and moving vehicles.
Citation Information
Patent Citations
Tracked robot intelligent reconnaissance system based on panoramic vision and multispectral analysis
CN120143291A
Ultra-wideband laser radar inertial navigation cooperative SLAM (Simultaneous Localization and Mapping) method and system
CN120403599A