Unmanned aerial vehicle aerial target positioning method and device based on road target scale features

CN122820847APending Publication Date: 2026-09-25SHANDONG JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611272895.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

综上所述,现有无人机对地目标定位技术方案存在以下不足:对GPS信号依赖程度高,在复杂环境中易受信号遮挡影响;基于视觉特征匹配的方法对场景纹理信息敏感,累计误差随时间增加;多传感器融合方案系统复杂、硬件成本高,难以轻量化部署;现有单目视觉方案缺乏基于目标尺度先验的异常筛选机制和多帧时序约束,无法有效抑制透视畸变和姿态抖动带来的定位误差

Benefits of technology

(一)无需外部依赖,自适应透视畸变校正

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820847A_ABST
    Figure CN122820847A_ABST
Patent Text Reader

Abstract

The application discloses a UAV aerial target positioning method and equipment based on road target scale features, and belongs to the technical field of image recognition and UAV aerial target positioning. The method is as follows: UAV aerial images are acquired and road traffic targets are recognized, geometric scale prior information of a corresponding category is matched in a pre-constructed road traffic target geometric size database according to the target category; a three-level progressive abnormality screening is adopted to output an effective target observation set; scale consistency constraints are constructed according to effective target detection frame information and geometric scale prior information, and a current frame perspective mapping matrix rough solution is solved; a time sequence sliding window is established to jointly iteratively optimize the perspective mapping matrix rough solution; global scale recovery is performed by using the optimized perspective mapping matrix, target image pixel positions are mapped to a world coordinate system, target longitude and latitude coordinates are calculated in combination with UAV longitude and latitude information, and geographic positioning is completed. The application can realize stable and accurate positioning of road traffic targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image recognition and UAV aerial target localization technology, and specifically relates to a method and device for UAV aerial target localization based on road target scale features. Background Technology

[0002] In recent years, with the continuous expansion of the highway network and the increasing number of road traffic infrastructures, the tasks of traffic facility inspection and road operation status monitoring have become increasingly arduous. Traditional road inspection mainly relies on manual on-site inspection and measurement, which is not only inefficient but also suffers from limited inspection scope, high operational risks, and long data update cycles, making it difficult to meet the actual needs of large-scale, routine, and intelligent road facility inspection. Spatial location information of targets such as vehicles illegally occupying lanes, road surface defects, damaged traffic ancillary facilities, and roadside obstacles is crucial foundational data for road maintenance, traffic operation management, emergency response, and intelligent transportation construction. Therefore, achieving rapid identification and accurate positioning of road traffic targets has significant engineering application value. Unmanned aerial vehicles (UAVs) have the advantages of high mobility, flexible deployment, wide inspection range, low cost, and non-contact operation, and have gradually become an important technical means for road traffic inspection. Utilizing UAVs to acquire road traffic scene images, combining them with target detection algorithms to complete road target identification, and further obtaining target spatial location information can provide reliable data support for road facility management, traffic incident monitoring, and digital road network construction, which is of great significance for improving the intelligence level of road traffic inspection. Currently, research on UAV visual perception technology mainly focuses on improving target detection accuracy, while research on target spatial positioning still has some shortcomings, especially in establishing the mapping relationship between image pixel coordinates and real geographic coordinates, calculating target spatial coordinates, and compensating for UAV pose errors. Existing methods still have considerable room for improvement.

[0003] Traditional UAV ground target localization primarily relies on the Global Positioning System (GPS), Inertial Navigation System (INS), and aerial photogrammetry. By combining the UAV's position, attitude, and camera parameters, the target pixel coordinates in the image are mapped to the geographic coordinate system to calculate the target's spatial location. This method is highly dependent on GPS signals and is easily affected by signal blockage and multipath propagation in complex environments such as urban canyons, under bridges, and indoors, leading to decreased positioning accuracy. To reduce reliance on satellite positioning signals, target localization methods based on visual odometry and visual simultaneous localization and mapping (VSLTD) have become widely used. These methods utilize feature extraction and matching between consecutive images to estimate the UAV's trajectory and combine multi-view geometric relationships to achieve ground target localization. They can achieve autonomous localization in GPS-free environments. However, they are highly dependent on scene texture information. When there are changes in lighting, motion blur, shadow occlusion, or weak texture areas, feature matching failures and increased cumulative errors can easily occur, affecting target localization accuracy. To further improve positioning performance, existing technologies employ information fusion from multiple sensors such as cameras, inertial measurement units, lidar, and GPS to enhance UAV pose estimation and target positioning accuracy. However, these solutions generally suffer from complex system composition, high hardware costs, insufficient real-time processing capabilities, and difficulties in multi-sensor time synchronization, limiting their application in lightweight UAV road traffic inspection scenarios. A search revealed that the closest existing technology to this application is patent CN202610027916.9 (UAV Ground Target Positioning Method and System Based on Monocular Vision). This solution establishes a multi-level coordinate transformation system, relying on real-time UAV pose, gimbal attitude data, and camera intrinsic parameters to achieve monocular ground target calculation. However, this solution entirely depends on fitting coordinate relationships using high-precision airborne hardware sensor parameters, lacks prior constraints on traffic target scale, cannot filter occlusion and false detection targets, and lacks a multi-frame temporal optimization mechanism, making positioning results prone to fluctuations due to UAV attitude jitter. In summary, existing UAV ground target positioning technologies have the following shortcomings: they are highly dependent on GPS signals and are easily affected by signal blockage in complex environments; visual feature matching methods are sensitive to scene texture information, and the cumulative error increases over time; multi-sensor fusion solutions are complex and costly, making lightweight deployment difficult; existing monocular vision solutions lack anomaly screening mechanisms based on target scale priors and multi-frame temporal constraints, and cannot effectively suppress positioning errors caused by perspective distortion and attitude jitter.

[0004] To address the aforementioned technical issues, there is an urgent need for a UAV aerial target localization method that can operate without relying on GPS signals and high-precision airborne sensing devices, without relying on scene texture feature matching, without requiring offline road calibration, and with target scale prior constraints and multi-frame temporal optimization capabilities. Summary of the Invention

[0005] To address the aforementioned technical challenges, this invention proposes a UAV aerial target localization method and device based on road target scale features. By constructing a priori database of road traffic target geometric scales, a three-level progressive anomaly screening process is employed to eliminate false detections, occlusions, and morphologically distorted targets. A normalized scale statistical model and a temporal sliding window are used to jointly optimize the perspective mapping matrix, adaptively correcting aerial perspective distortion and restoring the global scale. This achieves stable localization and geographic coordinate calculation of road traffic targets under monocular vision conditions.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, this invention proposes a UAV aerial target localization method based on road target scale features, comprising the following steps: Acquire aerial images taken by a drone, identify road traffic targets in the aerial images, and obtain category information, detection box location information, and detection confidence of each road traffic target; Based on the category information, the geometric scale prior information of the corresponding category is matched in the pre-constructed road traffic target geometric size database; based on the geometric scale prior information, a three-level progressive anomaly screening is performed on the road traffic targets, and a valid target observation set is output. Scale consistency constraints are constructed based on the correspondence between the detection box information of road traffic targets in the effective target observation set and the prior information of geometric scale, and the coarse solution of the perspective mapping matrix of the current frame is optimized based on the scale consistency constraints. Establish a temporal sliding window, and use the accumulated multi-frame data within the temporal sliding window to perform joint iterative optimization on the coarse solution of the perspective mapping matrix, and output the optimized perspective mapping matrix of the current frame. Global scale recovery is performed using the optimized perspective mapping matrix; based on the optimized perspective mapping matrix and the recovered global scale, the image pixel positions of the road traffic target are mapped to the world coordinate system to obtain the position coordinates of the road traffic target in the world coordinate system. Based on the location coordinates of the road traffic target in the world coordinate system and the latitude and longitude information of the UAV, the latitude and longitude coordinates of the road traffic target are calculated, and the geolocation is completed.

[0007] Furthermore, the road traffic target geometry database contains the true length, true width, true aspect ratio, and scale tolerance range of different categories of road traffic targets, and is stored in a structured manner using the target category number as an associated index. The prior information on geometric scale includes the true length, true width, true aspect ratio of the target, and scale tolerance range.

[0008] Furthermore, based on the aforementioned geometric scale prior information, a three-level progressive anomaly screening is performed on road traffic targets to output a valid target observation set, specifically: Road traffic targets with a detection confidence level lower than the preset confidence threshold are removed, completing the first-level screening; The remaining targets after the first-level screening are then processed. Based on the actual aspect ratio of the matched targets, the pixel aspect ratio of the detection box is calculated, and an aspect ratio tolerance coefficient is set. Road traffic targets with pixel aspect ratios exceeding the allowable range are then removed, thus completing the second-level screening. For the remaining targets after secondary screening, the mean and standard deviation of the pixel aspect ratio of the detection box are statistically analyzed by category to construct a judgment interval. Targets whose pixel aspect ratio exceeds the judgment interval are removed. After each removal, the judgment interval is re-statistically analyzed and updated based on the current remaining targets. This process is repeated until no targets are removed or the preset number of iterations is reached, thus completing the tertiary screening and outputting a set of valid target observations.

[0009] Furthermore, scale consistency constraints are constructed based on the correspondence between the detection box information of road traffic targets in the effective target observation set and the prior information of geometric scale, specifically as follows: The pixel length of each road traffic target is extracted using the minimum bounding rectangle of the target contour after perspective correction. The normalized scaling factor is calculated as follows: ; in, Indicates the pixel scale corresponding to a unit of actual length; This represents the average length of the category corresponding to the target obtained from database matching; The normalized scaling factor follows a normal distribution. For statistical hypotheses, where, This represents the normalized scale mean. Representing the normalized scale standard deviation, constructing a set of normalized scale factors. : ; in, This represents the total number of targets in the valid target observation set; Based on the assumption that the normalization scaling factors of each objective are independent, a joint likelihood function is constructed: ; in, Represents the set of normalized scaling factors The joint likelihood function value; Taking the negative logarithm of the joint likelihood function, a statistical loss function is constructed, wherein the statistical loss function... Represented as: ; in, Represents the perspective mapping matrix The statistical loss function value.

[0010] Furthermore, based on the aforementioned scale consistency constraint, the coarse solution of the perspective mapping matrix for the current frame is optimized, specifically as follows: A two-level optimization strategy combining maximum likelihood estimation and iterative weighted least squares is adopted, using the aforementioned statistical loss function. With the objective of minimizing the perspective mapping matrix, a coarse solution for the current frame's perspective mapping matrix is ​​iteratively obtained. The inner optimization involves fixing the current perspective mapping matrix and updating the mean and standard deviation of the normalized scale factor using the following formula: ; ; in, Indicates the first Normalized scalar mean at the next iteration; Indicates the first During the nth iteration Normalized scaling factor for each objective; Indicates the first Normalized scale standard deviation at the next iteration; The outer layer is optimized to have fixed updated statistical parameters. Calculate the residuals of each objective. ; in, Indicates the first The residuals of each objective; The optimization weights of each objective are dynamically adjusted based on the residual magnitude, and the perspective mapping matrix is ​​updated using an iterative weighted least squares method. Repeat the inner and outer layer optimizations until the convergence condition is met, and output the coarse solution of the perspective mapping matrix for the current frame. .

[0011] Furthermore, a temporal sliding window is established, and the coarse solution of the perspective mapping matrix is ​​jointly iteratively optimized using the accumulated multi-frame data within the temporal sliding window, outputting the optimized perspective mapping matrix for the current frame, specifically as follows: Establish a length of A time-series sliding window uses a first-in-first-out (FIFO) strategy to cache the most recently used data. The effective observation data; when a new frame enters the window, the earliest frame data in the window is removed to complete the window update; if the number of effective targets in the current frame is lower than the preset threshold, it is judged as a low-quality frame, not included in the window cache, and the optimization result of the previous frame is directly used. Construct a joint loss function that includes data items and a time-series smoothing term: ; in, This represents the value of the joint loss function. Indicates the time-series smoothing weighting coefficient; Represents a data item; Represents the time-series smoothing term; With joint loss function value With the goal of minimizing, the perspective mapping matrix of multiple frames within the window is jointly iteratively optimized. The inner optimization is to use the perspective mapping matrix of each frame within the current window to calculate the normalized scale factor of all effective targets, and then use maximum likelihood estimation to update the mean and standard deviation of the global normalized scale factor. The outer layer is optimized to have fixed updated global statistical parameters, using the joint loss function. With the goal of minimizing, the perspective mapping matrix of each frame within the window is updated simultaneously using an iterative weighted least squares method. The inner and outer optimizations are repeated until the convergence condition is met, and the optimized perspective mapping matrix of the current frame is output.

[0012] Furthermore, the convergence condition is as follows: convergence is determined when the total loss change is less than a first preset threshold, the iterative change of the current frame perspective mapping matrix is ​​less than a second preset threshold, and the difference between the perspective mapping matrices of adjacent frames is less than a third preset threshold.

[0013] Furthermore, global scale recovery is performed using the optimized perspective mapping matrix, specifically as follows: Extract the normalized scale mean obtained after the joint iterative optimization converges. Calculate the global scale factor : ; Using the global scale factor It converts pixel distances in the image domain into actual physical distances in the world domain, thus completing global scale recovery from pixel scale to physical scale.

[0014] Furthermore, based on the optimized perspective mapping matrix and the restored global scale, the image pixel positions of the road traffic targets are mapped to the world coordinate system to obtain the position coordinates of the road traffic targets in the world coordinate system, specifically: Establish the transformation relationships between the image pixel coordinate system, the image physical coordinate system, and the world coordinate system; The location of the detection box of the road traffic target in the image pixel coordinate system is converted into the image physical coordinates using camera intrinsic parameters; Under the assumption of orthophoto projection, the physical coordinates of the image are mapped to the world coordinate system using the optimized perspective mapping matrix to obtain the position coordinates of the road traffic target in the world coordinate system. Based on the global scale factor, the pixel distance in the image domain is converted into the actual physical distance in the world domain, thus completing the mapping of the target from the image pixel position to the position coordinates in the world coordinate system.

[0015] Secondly, this invention also proposes a UAV aerial target positioning device based on road target scale features, comprising: Image acquisition module The system includes at least one processor and a memory, the memory storing a computer program that, when executed by the at least one processor, implements the UAV aerial target localization method based on road target scale features.

[0016] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects: (i) Adaptive perspective distortion correction without external dependencies This invention employs a dual-constraint multi-frame joint optimization architecture of normalized scaling statistics and temporal smoothing. It does not rely on GPS signals, satellite maps, depth sensors, or high-precision pose sensing devices. It can dynamically solve the perspective mapping matrix using only a monocular aerial camera, adaptively correct aerial perspective distortion, and effectively suppress positioning jumps caused by UAV flight attitude jitter through sliding window temporal constraints. This significantly improves the temporal stability and scene adaptability of the positioning results. In contrast, existing optimal solutions are mostly fixed offline mapping models or single-frame geometric projection models. When the UAV attitude changes, the model fails and repeated calibration is required, resulting in poor temporal stability and scene adaptability.

[0017] (ii) Three-level progressive screening and robust optimization to ensure solution accuracy This invention employs a three-level progressive traffic target anomaly screening structure. Through scale database matching, confidence screening, geometric consistency verification, and iterative refinement, it progressively eliminates false detections, occlusions, and morphologically distorted anomalies, providing a high-quality and effective sample set for perspective calculation and fundamentally reducing the impact of anomalies on positioning accuracy. Simultaneously, this invention unifies the pixel scale of multiple target categories to the same physical scale benchmark by normalizing the scale factor, solving the interference problem of perspective correction caused by differences in target size. Furthermore, it utilizes a two-layer optimization strategy combining maximum likelihood estimation and iterative weighted least squares, adaptively adjusting the weights of each target based on the residual magnitude. This automatically reduces the interference of anomalies on perspective matrix calculation, maintaining high solution accuracy even in real-world scenarios with noise and outliers in target detection.

[0018] (iii) Lightweight deployment, outputting directly applicable location data This invention eliminates the need for offline road calibration or pre-built fixed mapping maps. The solution to the perspective mapping matrix relies entirely on traffic targets observed in the current aerial imagery. It adapts to changes in UAV flight altitude, viewpoint, and attitude, maintaining stable positioning performance in dynamic traffic scenarios. Furthermore, it requires only a monocular aerial camera, eliminating the need for expensive sensors such as LiDAR and high-precision IMUs, resulting in low computational resource consumption. It supports incremental update strategies to reduce real-time computation and can be deployed on lightweight UAV platforms, meeting the practical requirements of road traffic inspection for system cost and real-time performance. Finally, it restores the conversion relationship between pixel distance and actual physical distance through a global scale factor, outputting road traffic target positioning results with actual physical scale and precise latitude and longitude coordinates, providing reliable location data support for road infrastructure management, traffic incident monitoring, and digital road network construction. Attached Figure Description

[0019] Figure 1 This is a flowchart of the UAV aerial target localization method based on road target scale features proposed in Embodiment 1 of the present invention; Figure 2 This is an architecture diagram of the UAV aerial target localization method based on road target scale features proposed in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the coordinate transformation relationship proposed in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the UAV aerial target positioning device based on road target scale features proposed in Embodiment 2 of the present invention. Detailed Implementation

[0020] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components, processing techniques, and processes are omitted in this invention to avoid unnecessarily limiting the invention.

[0021] Example 1 Embodiment 1 of this invention proposes a UAV aerial photography target localization method based on road target scale features. This method addresses the problems of existing UAV aerial photography target localization methods, such as strong dependence on GPS signals, low positioning accuracy in complex environments, easy failure of scene texture feature matching, complex and costly multi-sensor fusion schemes, lack of prior constraints on traffic target scale and multi-frame temporal optimization mechanisms, and inability to achieve stable and accurate positioning of road traffic targets under lightweight deployment conditions.

[0022] The implementation process of Embodiment 1 of this invention can be summarized as follows: A monocular aerial camera is mounted on a UAV aerial photography target localization device based on road target scale features, and a road traffic target geometric scale database is built-in. Road traffic targets are identified through a target detection algorithm. Based on the target category, corresponding standard geometric size information is matched. After confidence screening, target aspect ratio geometric verification, and iterative refinement, falsely detected targets, occluded targets, and targets with abnormal shapes are eliminated. Based on the retained valid targets, a normalized scale factor is constructed. The MLE-IRLS algorithm (Maximum Likelihood Estimation - Iteratively Reweighted Least Squares, a two-layer optimization strategy combining maximum likelihood estimation and iterative weighted least squares) is used to solve the single-frame perspective mapping matrix. Combined with a temporal sliding window to cache multiple frames of valid observation data, a joint optimization model including scale fitting and temporal smoothing terms is established to jointly optimize the perspective mapping matrix across multiple frames, achieving dynamic correction of perspective distortion in the aerial image. The optimized perspective mapping matrix is ​​used to complete the transformation from image pixel coordinates to ground world coordinates, obtaining the geographical location information of the road traffic targets. This invention does not rely on satellite maps, depth sensors, or offline road calibration; it can achieve stable positioning of road traffic targets in dynamic traffic scenarios using only monocular vision.

[0023] Figure 1 This is a flowchart of the UAV aerial target localization method based on road target scale features proposed in Embodiment 1 of the present invention; Figure 2 This is an architecture diagram of the UAV aerial target localization method based on road target scale features proposed in Embodiment 1 of the present invention; combined with Figure 1 and Figure 2 The process of implementing this invention will be explained together.

[0024] In step S1: acquire aerial images taken by the UAV, identify road traffic targets in the aerial images, and acquire category information, detection box location information, and detection confidence of each road traffic target; In this invention, road traffic target detection is used to identify road traffic targets in drone aerial images.

[0025] The drone is equipped with an optical camera to collect images of road traffic scenes. Then, a target detection algorithm is used to identify vehicle targets and road traffic structures in the images, and obtain information such as target category, detection box position and detection confidence.

[0026] This invention employs the YOLO series of target detection algorithms to identify road traffic targets in drone aerial images. The YOLO (You Only Look Once) algorithm, a representative method of single-stage target detection, transforms the target detection task into an end-to-end regression problem. It directly predicts the class probability and bounding box coordinates of road traffic targets from drone aerial images using a single neural network. It features high detection speed and high accuracy, making it suitable for applications with limited computing resources on drone platforms that require real-time processing.

[0027] Table 1: Target Labels

[0028] This invention classifies road traffic targets and establishes a target detection and labeling system. Vehicle targets include minibuses, small passenger cars, medium-sized passenger cars, large passenger cars, as well as mini trucks, light trucks, medium trucks, and heavy trucks. Road traffic structures include lane lines, traffic signs, and directional arrows.

[0029] Based on the established labeling system, target detection algorithms such as YOLO are used to detect vehicle targets in road traffic scene images collected by UAVs. and road traffic structures The test results are expressed as follows: ; ; in, Represented as the first The coordinates of the top left corner of the vehicle detection frame. Represented as the first The coordinates of the lower right corner of the vehicle detection frame. Represented as the first The coordinates of the top left corner of the structure inspection frame. Represented as the first The coordinates of the lower right corner of the structure detection frame. , Separate vehicle target and structure target types, , This indicates the confidence level of vehicle and structure target detection.

[0030] In step S2, based on the category information, the geometric scale prior information of the corresponding category is matched in the pre-constructed road traffic target geometric size database; based on the geometric scale prior information, a three-level progressive anomaly screening is performed on the road traffic targets, and a valid target observation set is output. S2.1: This invention requires obtaining the true geometric scale information of road traffic targets and constructing a road traffic target geometric dimension database. The road traffic target geometric dimension database contains the true length, true width, true aspect ratio, and scale tolerance range of different categories of road traffic targets, and is structured and stored using the target category number as an associated index; The true scale information of road traffic targets is defined as: ; in, Indicates the average length of road traffic targets; Indicates the average width of the road traffic target; Indicates the upper limit of road traffic targets; Indicates the lower limit of the road traffic target; This represents the minimum target value for road traffic. Indicates the maximum target value for road traffic; The target is the true aspect ratio.

[0031] S2.2: The database adopts a structured storage method to uniformly manage target category information and geometric dimension information, and uses the target category number as the association index to realize the association between different data.

[0032] The target category data structure is used to store road traffic target category information and establish the correspondence between target categories and YOLO detection labels.

[0033] Target category data structure Defined as: ; Where Class_ID is the target category number, Class_Name is the target category name, and Label_Name is the target detection label name.

[0034] The target category data structure is used for target category index matching and database retrieval.

[0035] The target geometric scale data structure is used to store the prior geometric scale information of road traffic targets, including the target's actual physical scale information and scale constraint information.

[0036] Target geometric scale data structure Defined as: ; in, Indicates the geometric scale number.

[0037] The target geometric scale data structure is mainly used for target true scale matching, scale constraint verification, and abnormal target screening.

[0038] The database uses the target category number (Class_ID) as a unified association index to achieve data association between target category data and target geometric scale data. This unified category index enables rapid retrieval and scale matching of target scale information.

[0039] The database uses a structured table format to store data, with each target category corresponding to a unique category number and each scale data point corresponding to a unique geometric scale number. This structured data organization method enables standardized management and dynamic updates of target scale data.

[0040] S2.3: Based on the prior information of the geometric scale, a three-level progressive anomaly screening is performed on road traffic targets to output a valid target observation set. Specifically, road traffic targets with a detection confidence level lower than a preset confidence threshold are removed, completing the first-level screening; for the remaining targets after the first-level screening, the pixel aspect ratio of the detection box is calculated based on the true aspect ratio of the matched targets, and an aspect ratio tolerance coefficient is set. Road traffic targets with pixel aspect ratios exceeding the allowable range are removed, completing the second-level screening; for the remaining targets after the second-level screening, the mean and standard deviation of the pixel aspect ratio of the detection box are statistically analyzed by category to construct a judgment interval. Targets with pixel aspect ratios exceeding the judgment interval are removed; after each removal, the judgment interval is re-statistically updated based on the current remaining targets. This process is repeated until no targets are removed or the preset number of iterations is reached, completing the third-level screening and outputting a valid target observation set.

[0041] Specifically, the target detection results of this invention contain false detections, occlusions, and detection box offsets, leading to distortion of some target geometric features. To improve the reliability of the detection results, this invention utilizes a road traffic target geometric dimension database to construct geometric scale constraints, performs anomaly screening on the detected targets, removes targets that do not meet the geometric constraints, and obtains a valid target observation set.

[0042] First, target category labels are obtained based on the target detection results. Using the target category as a retrieval index, the corresponding category's geometric dimension information is queried in the road traffic target geometry dimension database, and a correspondence between the detected target and the actual geometric dimension is established.

[0043] After completing database matching, construct a scale-matching target set. : ; in, For the first The detection bounding box information for each target (including vehicles and buildings). This provides the true geometric scale information for the category corresponding to the target. This represents the initial total number of targets to be detected.

[0044] Due to issues such as occlusion, motion blur, and complex backgrounds in drone aerial images, some target detection results exhibit false positives and false negatives.

[0045] Therefore, the effectiveness of the detection targets is first screened based on the target detection confidence level.

[0046] Set the reliability threshold to When the following conditions are met: If the target is detected correctly, it is considered a valid target; otherwise, it is considered an invalid target and removed. Confidence-based screening can reduce the impact of false positives on subsequent scale recovery.

[0047] After confidence level screening, a set of credible targets is established. : ; in, This represents the number of targets that pass the first level (confidence level) screening, i.e., the set of credible targets. The total number of targets retained.

[0048] For the remaining targets after confidence screening, the shape of the detection box is validated for rationality based on the standard aspect ratio constraints of the target geometry database, and abnormal samples with severe truncation, large-scale occlusion, and morphological distortion are further eliminated.

[0049] Based on the database matching results, calculate the pixel aspect ratio of the original detection box of the target. :

[0050] in, , These are the pixel width and pixel height of the detection box, respectively; Considering the morphological deviations caused by perspective distortion and detection frame positioning errors, an aspect ratio tolerance coefficient is set. Define a reasonable range for the form: ; When the aspect ratio of a target pixel exceeds the allowed range, it is determined that the target has abnormalities such as occlusion, truncation, or detection box offset, and it is removed from the target set.

[0051] To reduce the impact of outliers on the statistical results, an iterative approach is used to update the screening results. After each outlier removal, the pixel aspect ratio distribution of the remaining samples is recalculated according to the target category, and the geometric consistency judgment interval for the corresponding category is updated. The outlier judgment process is repeated until no new outliers are found, or the preset number of iterations is reached.

[0052] The final output is the cleaned set of valid target observations: ; in, This set represents the effective number of targets and will be used for subsequent calculations of the perspective mapping matrix and scale recovery.

[0053] In step S3, a scale consistency constraint is constructed based on the correspondence between the detection box information of road traffic targets in the effective target observation set and the prior information of geometric scale, and the coarse solution of the perspective mapping matrix of the current frame is optimized based on the scale consistency constraint. Although the valid target observation set obtained in the previous step eliminated abnormal targets, the single-frame perspective mapping matrix may still have solution errors due to factors such as changes in UAV flight attitude, perspective distortion, and target detection errors. To improve the stability of the perspective matrix solution and the accuracy of scale recovery, this invention first constructs a normalized scale statistical model using the actual geometric dimensions of road traffic targets to solve the single-frame perspective mapping matrix. Based on this, a temporal sliding window is introduced to jointly optimize the perspective mapping matrices of multiple consecutive frames, achieving stable estimation of the perspective matrix and global scale recovery.

[0054] In S3.1, scale consistency constraints are constructed based on the correspondence between the detection bounding box information of road traffic targets in the effective target observation set and the prior information of geometric scale. Specifically: The pixel length of each road traffic target is extracted using the minimum bounding rectangle of the target contour after perspective correction. Since different target categories have different true sizes, pixel length cannot be used for direct comparison. To eliminate the influence of differences in true size between different vehicle models, a normalized scaling factor is introduced: ; in, Indicates the pixel scale corresponding to a unit of actual length; This represents the average length of the category corresponding to the target obtained from database matching; In theory, once perspective distortion is correctly eliminated, the pixel scale per unit length corresponding to different target categories should tend to be consistent. Considering the combined influence of various random factors such as target detection error, size fluctuation, and perspective transformation residuals, it is assumed that the normalized scale factor approximately follows a normal distribution. For statistical hypotheses, where, This represents the normalized scale mean. Representing the normalized scale standard deviation, constructing a set of normalized scale factors. : ; in, This represents the total number of targets in the valid target observation set; To find the optimal perspective mapping matrix, a statistical model is established using the maximum likelihood estimation approach. Given a perspective mapping matrix... Under the condition of [condition], and based on the assumption that the normalization scaling factors of each objective are mutually independent, a joint likelihood function is constructed: ; in, Represents the set of normalized scaling factors The joint likelihood function value; Taking the negative logarithm of the joint likelihood function, a statistical loss function is constructed, wherein the statistical loss function... Represented as: ; in, Represents the perspective mapping matrix The statistical loss function value. This loss function reflects the degree to which the normalized scaling factor deviates from the overall statistical distribution. The more concentrated the normalized scaling factors are for different objectives, the better. The smaller the value, the smaller the loss function value, indicating that the perspective correction result is closer to the real bird's-eye view.

[0055] In S3.2, a two-layer iterative optimization strategy of MLE and IRLS is adopted to improve the accuracy of solving the perspective matrix.

[0056] A two-level optimization strategy combining maximum likelihood estimation and iterative weighted least squares is adopted, using the aforementioned statistical loss function. With the objective of minimizing the perspective mapping matrix, a coarse solution for the current frame's perspective mapping matrix is ​​iteratively obtained. The inner optimization involves fixing the current perspective mapping matrix and updating the mean and standard deviation of the normalized scale factor using the following formula: ; ; in, Indicates the first Normalized scalar mean at the next iteration; Indicates the first During the nth iteration The normalized scaling factor for each objective; Indicates the first Normalized scale standard deviation at the next iteration; The outer layer is optimized to have fixed updated statistical parameters. Calculate the residuals of each objective. ; in, Indicates the first The residuals of each objective; The optimization weights of each objective are dynamically adjusted based on the residual magnitude. For objectives with smaller residuals, their optimization weights are increased; for abnormal objectives with larger residuals due to occlusion, false detections, or bounding box offsets, their optimization weights are automatically decreased.

[0057] Subsequently, the perspective mapping matrix is ​​updated using an iterative weighted least squares method. Repeat the inner and outer layer optimizations until the convergence condition is met, and output the coarse solution of the perspective mapping matrix for the current frame. .

[0058] The optimization objective is to minimize the statistical loss function. ; in, This is represented as the coarse solution of the optimal perspective mapping matrix obtained from a single frame view. The algorithm is considered to have converged when the following conditions are met: the iteration stops when both the change in loss and the change in the elements of the perspective matrix are less than a preset threshold, and a single-frame perspective coarse solution is output.

[0059] In step S4, a temporal sliding window is established, and the coarse solution of the perspective mapping matrix is ​​jointly iteratively optimized using the multi-frame data accumulated within the temporal sliding window, and the optimized perspective mapping matrix of the current frame is output. S4.1: To address the issue that single-frame solutions are susceptible to drone attitude jitter and sparse scene targets, resulting in large frame-by-frame fluctuations and insufficient stability of the perspective matrix, a temporal sliding window frame buffering mechanism is introduced on the basis of single-frame solutions. Through joint optimization of multi-frame data and temporal smoothing constraints, the solution accuracy and temporal consistency of the perspective matrix are further improved.

[0060] Define the length as A time-series sliding window uses a first-in-first-out (FIFO) strategy to cache the most recently used data. The effective observation data for each frame. The data cached for each frame includes: frame number, the set of effective target observations after anomaly filtering, the perspective mapping matrix obtained from coarse optimization of that frame, and the inter-frame camera pose transformation matrix.

[0061] When a new frame enters the window, the oldest frame in the window is removed, and the window is updated.

[0062] If the number of valid targets in the current frame is lower than the preset threshold, it is judged as a low-quality frame and is not included in the window cache. Instead, the optimization result of the previous frame is used directly to avoid degrading the global statistical effect.

[0063] Based on the single-frame maximum likelihood statistical loss, a joint loss function is constructed that includes data terms and temporal smoothing terms, balancing scaling accuracy and temporal stability. ; in, This represents the value of the joint loss function. Indicates the time-series smoothing weighting coefficient; Represents a data item; Represents the time-series smoothing term; Data Items This is the sum of the normalized scaled negative log-likelihoods of all valid targets across all frames within the sliding window. It continues the statistical constraint logic of single-frame maximum likelihood estimation, and improves statistical reliability by expanding the sample size. ; in, The frame number of the current frame to be optimized. For the first Number of valid targets per frame For the aligned first Frame number One target normalized scaling factor.

[0064] Time series smoothing term Constraining the differences in perspective mapping matrices between adjacent frames suppresses perspective result jumps caused by high-frequency jitter in UAV attitude, thereby improving the temporal stability of the output. ; in, Let Frobenius norm be the matrix. For the first The perspective mapping matrix of a frame.

[0065] S4.2, using the joint loss function value With the goal of minimizing, the perspective mapping matrix of multiple frames within the window is jointly iteratively optimized. The inner optimization is to use the perspective mapping matrix of each frame within the current window to calculate the normalized scale factor of all effective targets, and then use maximum likelihood estimation to update the mean and standard deviation of the global normalized scale factor. Global normalized scale mean update formula: ; Global normalized scale standard deviation update formula: .

[0066] The outer layer is optimized to have fixed updated global statistical parameters, using the joint loss function. With the goal of minimizing, the perspective mapping matrix of each frame within the window is updated simultaneously using an iterative weighted least squares method. The inner and outer optimizations are repeated until the convergence condition is met, and the optimized perspective mapping matrix of the current frame is output.

[0067] To reduce computational load and meet real-time requirements, an incremental update strategy is adopted: when a new frame enters the sliding window, only the data constraints of the oldest frame are removed, and convergence can be achieved by performing a small number of iterations based on the optimization results of the previous round, without having to re-solve the entire window of data.

[0068] The algorithm is considered convergent if the following conditions are met: Changes in total losses: ; Current frame matrix iteration changes: ; Differences in the matrix between adjacent frames: ; in, This represents the joint loss function value calculated in the current iteration. This represents the joint loss function value obtained from the previous iteration. This indicates the first preset threshold. Indicates the number of iterations in the current iteration round. The perspective mapping matrix after frame update; Indicates the th iteration in the previous iteration frame; This indicates the second preset threshold. Indicates the number of iterations in the current iteration round. The perspective mapping matrix of the previous frame; This indicates the third preset threshold.

[0069] In step S5, the optimized perspective mapping matrix is ​​used to perform global scale recovery; based on the optimized perspective mapping matrix and the recovered global scale, the image pixel positions of the road traffic target are mapped to the world coordinate system to obtain the position coordinates of the road traffic target in the world coordinate system. The optimized perspective mapping matrix H* is used to perform perspective transformation on the original UAV image to generate a road scene image with a unified bird's-eye view.

[0070] Extract the normalized scale mean obtained after the joint iterative optimization converges. ,because This represents the average pixel scale (pixel / m) corresponding to a unit of actual length; therefore, its reciprocal is the actual physical length (m / pixel) corresponding to a unit of pixel. Calculate the global scale factor. : ; Using the global scale factor It converts pixel distances in the image domain into actual physical distances in the world domain, thus completing global scale recovery from pixel scale to physical scale.

[0071] Figure 3 This is a schematic diagram of the coordinate transformation relationship proposed in Embodiment 1 of the present invention; Establish the transformation relationships between the image pixel coordinate system, the image physical coordinate system, and the world coordinate system.

[0072] Image pixel coordinate system: Taking the top left corner of the image as the origin, shaft and The axes are respectively related to the image coordinate system. shaft and The axes are parallel, and the unit is... .

[0073] Image physical coordinate system: The optical center is the midpoint of the image. The axis is parallel to the horizontal right of the image. Axis perpendicular to Axial downwards, unit is .

[0074] World coordinate system: This represents the object's position coordinates in actual space. The origin is the point where the drone's vertical projection onto the ground is located. The axis is aligned with the drone's flight direction. shaft and parallel, The axis points upwards, and the unit is... .

[0075] Convert pixel coordinates to physical coordinates: ; in, The pixel x-coordinate of the image center point; The ordinate of the pixel representing the center point of the image; Represents the global scale factor; This represents the x-coordinate of the target in the image's physical coordinate system; This represents the vertical coordinate of the target in the image's physical coordinate system.

[0076] Under the assumption of orthophoto projection, the physical coordinates of the image are mapped to the world coordinate system: ; Distance between the drone and the vehicle : ; in, This represents the x-coordinate of the target in the world coordinate system; This represents the vertical coordinate of the target in the world coordinate system.

[0077] In step S6, the latitude and longitude coordinates of the road traffic target are calculated based on the position coordinates of the road traffic target in the world coordinate system and the latitude and longitude information of the UAV, thus completing the geolocation.

[0078] Convert the target's position offset in the world coordinate system into geographic coordinate offset.

[0079] The mathematical model for latitude and longitude offset is: Ground distance corresponding to each degree of latitude:

[0080] Ground distance corresponding to each degree of longitude:

[0081] in, The latitude (in radians) of the drone.

[0082] The latitude and longitude offset of the vehicle relative to the drone is: ; in, This is the latitude offset. This represents the longitude offset.

[0083] The final calculation process for the vehicle's latitude and longitude is as follows: Let the latitude and longitude of the UAV be... Then the vehicle's latitude and longitude for: .

[0084] The above calculations complete the geographic location of the road traffic target.

[0085] The UAV aerial photography target localization method based on road target scale features proposed in Embodiment 1 of this invention provides a stable scale reference benchmark for target localization, unaffected by ambient lighting and texture, by constructing a prior database of road traffic target geometric scales. A three-level progressive anomaly screening mechanism effectively eliminates false detections, occlusions, and morphological distortions, ensuring the reliability of target samples participating in perspective calculation. A normalized scale statistical model unifies multiple target categories to the same physical scale benchmark, resolving the interference of different vehicle size differences on perspective correction. A two-layer MLE-IRLS optimization strategy adaptively adjusts weights based on residuals, further reducing the impact of anomaly samples on perspective matrix calculation. Finally, a temporal sliding window multi-frame joint optimization is introduced, effectively suppressing localization jumps caused by UAV attitude jitter by utilizing temporal consistency constraints between adjacent frames, significantly improving the temporal stability and scene adaptability of the localization results. The entire method flow is interconnected, from target detection, anomaly screening, single-frame coarse solution to multi-frame refinement, forming a complete UAV aerial photography target localization solution that requires no external dependencies.

[0086] Example 2 Based on the UAV aerial photography target localization method based on road target scale features proposed in Embodiment 1 of the present invention, Embodiment 2 of the present invention also proposes a UAV aerial photography target localization device based on road target scale features. Figure 4 This is a schematic diagram of the UAV aerial target positioning device based on road target scale features proposed in Embodiment 2 of the present invention. At the hardware level, the electronic device 400 includes a processor 410, and optionally, an internal bus 420, a network interface 430, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or it may also include non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations. The processor 410, network interface 430, and memory can be interconnected via an internal bus 420. This internal bus 420 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be categorized as an address bus, data bus, control bus, etc. For ease of illustration, only a single bidirectional arrow is used in this diagram, but this does not imply that there is only one bus or one type of bus. The memory is used to store programs. Specifically, the program can include program code, which includes computer operation instructions. The memory can include main memory 440 and non-volatile memory 450, and provides instructions and data to the processor 410. The processor 410 reads the corresponding computer program from the non-volatile memory 450 into the main memory 440 and then runs it, forming a UAV aerial photography target positioning device at the logical level. The processor 410 executes the program stored in the memory and specifically performs the following: Step S1: Acquire aerial images taken by the UAV, identify road traffic targets in the aerial images, and obtain the category information, detection box location information, and detection confidence of each road traffic target; Step S2: Based on the category information, match the geometric scale prior information of the corresponding category in the pre-constructed road traffic target geometric size database; based on the geometric scale prior information, perform a three-level progressive anomaly screening of road traffic targets and output a set of valid target observations; Step S3: Construct scale consistency constraints based on the correspondence between the detection box information of road traffic targets in the effective target observation set and the prior information of geometric scale, and optimize the coarse solution of the perspective mapping matrix of the current frame based on the scale consistency constraints; Step S4: Establish a temporal sliding window, use the accumulated multi-frame data within the temporal sliding window to perform joint iterative optimization on the coarse solution of the perspective mapping matrix, and output the optimized perspective mapping matrix of the current frame; Step S5: Perform global scale recovery using the optimized perspective mapping matrix; based on the optimized perspective mapping matrix and the recovered global scale, map the image pixel positions of the road traffic target to the world coordinate system to obtain the position coordinates of the road traffic target in the world coordinate system; Step S6: Based on the location coordinates of the road traffic target in the world coordinate system and the latitude and longitude information of the UAV, calculate the latitude and longitude coordinates of the road traffic target to complete the geolocation.

[0087] Figure 1 It can be applied to processor 410, or implemented by processor 410. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The processor mentioned above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0088] The description of the relevant parts of the UAV aerial target positioning device based on road target scale features provided in Embodiment 2 of this application can be found in the detailed description of the corresponding parts of the UAV aerial target positioning method based on road target scale features provided in Embodiment 1 of this application, and will not be repeated here.

[0089] The UAV aerial photography target positioning device based on road target scale features provided in Embodiment 2 of the present invention can achieve the same technical effect as the UAV aerial photography target positioning method based on road target scale features in Embodiment 1 of the present invention.

[0090] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that the elements inherent in a process, method, article, or apparatus that includes a list of elements are included. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Additionally, portions of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.

[0091] While specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art can make other modifications or variations based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for locating targets in UAV aerial photography based on road target scale features, characterized in that, Includes the following steps: Acquire aerial images taken by a drone, identify road traffic targets in the aerial images, and obtain category information, detection box location information, and detection confidence of each road traffic target; Based on the category information, the geometric scale prior information of the corresponding category is matched in the pre-constructed road traffic target geometric size database; based on the geometric scale prior information, a three-level progressive anomaly screening is performed on the road traffic targets, and a valid target observation set is output. Scale consistency constraints are constructed based on the correspondence between the detection box information of road traffic targets in the effective target observation set and the prior information of geometric scale, and the coarse solution of the perspective mapping matrix of the current frame is optimized based on the scale consistency constraints. A temporal sliding window is established, and the coarse solution of the perspective mapping matrix is ​​jointly iteratively optimized using the multi-frame data accumulated within the temporal sliding window. The optimized perspective mapping matrix of the current frame is then output. Global scale recovery is performed using the optimized perspective mapping matrix; Based on the optimized perspective mapping matrix and the restored global scale, the image pixel positions of the road traffic target are mapped to the world coordinate system to obtain the position coordinates of the road traffic target in the world coordinate system. Based on the location coordinates of the road traffic target in the world coordinate system and the latitude and longitude information of the UAV, the latitude and longitude coordinates of the road traffic target are calculated, and the geolocation is completed.

2. The method according to claim 1, characterized in that, The road traffic target geometry database contains the actual length, actual width, actual aspect ratio, and scale tolerance range of different categories of road traffic targets, and is stored in a structured manner using the target category number as an associated index. The prior information on geometric scale includes the true length, true width, true aspect ratio of the target, and scale tolerance range.

3. The method according to claim 2, characterized in that, Based on the aforementioned geometric scale prior information, a three-level progressive anomaly screening is performed on road traffic targets to output a set of valid target observations, specifically: Road traffic targets with a detection confidence level lower than the preset confidence threshold are removed, completing the first-level screening; The remaining targets after the first-level screening are then processed. Based on the actual aspect ratio of the matched targets, the pixel aspect ratio of the detection box is calculated, and an aspect ratio tolerance coefficient is set. Road traffic targets with pixel aspect ratios exceeding the allowable range are then removed, thus completing the second-level screening. For the remaining targets after secondary screening, the mean and standard deviation of the pixel aspect ratio of the detection box are statistically analyzed by category to construct a judgment interval, and targets whose pixel aspect ratio exceeds the judgment interval are removed. After each elimination, the judgment interval is re-statistically updated based on the remaining targets. This process is repeated until no targets are eliminated or the preset number of iterations is reached, completing the three-level screening and outputting a set of valid target observations.

4. The method according to claim 1, characterized in that, Scale consistency constraints are constructed based on the correspondence between the detection bounding box information and the geometric scale prior information of road traffic targets in the effective target observation set, specifically as follows: The pixel length of each road traffic target is extracted using the minimum bounding rectangle of the target contour after perspective correction. The normalized scaling factor is calculated as follows: ; in, Indicates the pixel scale corresponding to a unit of actual length; This represents the average length of the category corresponding to the target obtained from database matching; The normalized scaling factor follows a normal distribution. For statistical hypotheses, where, This represents the normalized scale mean. Representing the normalized scale standard deviation, constructing a set of normalized scale factors. : ; in, This represents the total number of targets in the valid target observation set; Based on the assumption that the normalization scaling factors of each objective are independent, a joint likelihood function is constructed: ; in, Represents the set of normalized scaling factors The joint likelihood function value; Taking the negative logarithm of the joint likelihood function, a statistical loss function is constructed, wherein the statistical loss function... Represented as: ; in, Represents the perspective mapping matrix The statistical loss function value.

5. The method according to claim 4, characterized in that, The coarse solution of the perspective mapping matrix for the current frame is optimized based on the aforementioned scale consistency constraint, specifically as follows: A two-level optimization strategy combining maximum likelihood estimation and iterative weighted least squares is adopted, using the aforementioned statistical loss function. With the objective of minimizing the perspective mapping matrix, a coarse solution for the current frame's perspective mapping matrix is ​​iteratively obtained. The inner optimization involves fixing the current perspective mapping matrix and updating the mean and standard deviation of the normalized scale factor using the following formula: ; ; in, Indicates the first Normalized scalar mean at the next iteration; Indicates the first During the nth iteration Normalized scaling factor for each objective; Indicates the first Normalized scale standard deviation at the next iteration; The outer layer is optimized to have fixed updated statistical parameters. Calculate the residuals of each objective. ; in, Indicates the first The residuals of each objective; The optimization weights of each objective are dynamically adjusted based on the residual magnitude, and the perspective mapping matrix is ​​updated using an iterative weighted least squares method. Repeat the inner and outer layer optimizations until the convergence condition is met, and output the coarse solution of the perspective mapping matrix for the current frame. .

6. The method according to claim 5, characterized in that, A temporal sliding window is established, and the coarse solution of the perspective mapping matrix is ​​jointly iteratively optimized using multi-frame data accumulated within the temporal sliding window. The optimized perspective mapping matrix for the current frame is then output, specifically as follows: Establish a length of A time-series sliding window uses a first-in-first-out (FIFO) strategy to cache the most recently used data. The effective observation data; when a new frame enters the window, the earliest frame data in the window is removed to complete the window update; if the number of effective targets in the current frame is lower than the preset threshold, it is judged as a low-quality frame, not included in the window cache, and the optimization result of the previous frame is directly used. Construct a joint loss function that includes data items and a time-series smoothing term: ; in, This represents the value of the joint loss function. Indicates the time-series smoothing weighting coefficient; Represents a data item; Represents the time-series smoothing term; With joint loss function value With the goal of minimizing, the perspective mapping matrix of multiple frames within the window is jointly iteratively optimized. The inner optimization is to use the perspective mapping matrix of each frame within the current window to calculate the normalized scale factor of all effective targets, and then use maximum likelihood estimation to update the mean and standard deviation of the global normalized scale factor. The outer layer is optimized to have fixed updated global statistical parameters, using the joint loss function. With the goal of minimizing, the perspective mapping matrix of each frame within the window is updated simultaneously using an iterative weighted least squares method. The inner and outer optimizations are repeated until the convergence condition is met, and the optimized perspective mapping matrix of the current frame is output.

7. The method according to claim 6, characterized in that, The convergence condition is as follows: convergence is determined when the total loss change is less than a first preset threshold, the iterative change of the current frame perspective mapping matrix is ​​less than a second preset threshold, and the difference between the perspective mapping matrices of adjacent frames is less than a third preset threshold.

8. The method according to claim 1, characterized in that, Global scale recovery is performed using the optimized perspective mapping matrix, specifically as follows: Extract the normalized scale mean obtained after the joint iterative optimization converges. Calculate the global scale factor : ; Using the global scale factor It converts pixel distances in the image domain into actual physical distances in the world domain, thus completing global scale recovery from pixel scale to physical scale.

9. The method according to claim 1, characterized in that, Based on the optimized perspective mapping matrix and the restored global scale, the image pixel positions of the road traffic targets are mapped to the world coordinate system to obtain the position coordinates of the road traffic targets in the world coordinate system, specifically: Establish the transformation relationships between the image pixel coordinate system, the image physical coordinate system, and the world coordinate system; The location of the detection box of the road traffic target in the image pixel coordinate system is converted into the image physical coordinates using camera intrinsic parameters; Under the assumption of orthophoto projection, the physical coordinates of the image are mapped to the world coordinate system using the optimized perspective mapping matrix to obtain the position coordinates of the road traffic target in the world coordinate system. Based on the global scale factor, the pixel distance in the image domain is converted into the actual physical distance in the world domain, thus completing the mapping of the target from the image pixel position to the position coordinates in the world coordinate system.

10. A drone-based aerial target localization device based on road target scale features, including: Image acquisition module At least one processor and a memory, the memory storing a computer program, characterized in that, When the computer program is executed by the at least one processor, it implements the UAV aerial target localization method based on road target scale features as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Unmanned aerial vehicle ground target positioning method and system based on monocular vision

    CN121962289A