An unmanned aerial vehicle visual radar close-coupling positioning method, device, equipment, medium and product

By employing semantic segmentation and tightly coupled filtering optimization methods, the problems of dynamic interference, mismatch, and loose fusion in UAV SLAM systems were solved, achieving high-precision and robust visual radar tightly coupled localization, generating structured semantic maps, and enhancing the autonomous navigation capabilities of UAVs.

CN121432459BActive Publication Date: 2026-05-08NORTH CHINA UNIVERSITY OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTH CHINA UNIVERSITY OF TECHNOLOGY
Filing Date
2025-12-02
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing UAV SLAM systems have shortcomings in dynamic interference suppression, feature matching reliability, multi-source data fusion, and map building practicality, resulting in insufficient positioning accuracy and robustness, especially in complex dynamic and structured scenarios.

Method used

Image data is processed using the YOLOv8-seg semantic segmentation network. Feature points are selected based on the semantic consistency principle, and a structural semantic plane is extracted through semantic-geometric consistency to construct a semantic map. Tightly coupled filtering optimization is performed using IESKF to achieve tight-coupled positioning of visual radar.

Benefits of technology

It improves the accuracy and robustness of UAV positioning, reduces the impact of dynamic interference and mismatch, generates a structured semantic map with practical application value, and enhances the system's autonomous navigation capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121432459B_ABST
    Figure CN121432459B_ABST
Patent Text Reader

Abstract

The application discloses a method, device, equipment, medium and product for close-coupled positioning of a visual radar of a UAV, and relates to the technical field of autonomous navigation of a UAV. The method comprises: processing image data by using a semantic segmentation network, screening feature points according to a semantic segmentation result, and performing feature matching constraint on the screened static feature points based on a semantic consistency principle; projecting point cloud data to an image coordinate system, giving semantic information, and performing semantic filtering processing; performing plane feature extraction on structural semantic point cloud based on semantic-geometric consistency as a constraint condition, merging and optimizing a plurality of semantic planes, and obtaining a structural semantic map of an environment; determining an observation model based on a point-plane constraint residual and a semantic weight coefficient, and performing close-coupled filtering and optimization processing by using an IESKF to obtain updated state estimation information. The application can improve the positioning accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous navigation technology for unmanned aerial vehicles (UAVs), and in particular to a method, apparatus, equipment, medium, and product for UAV visual radar tightly coupled positioning. Background Technology

[0002] With the development of the low-altitude economy, real-time accurate Simultaneous Localization and Mapping (SLAM) technology is the core foundation for achieving full autonomy in applications such as UAV autonomous navigation, UAV inspection, and smart warehousing. Visual sensors provide rich environmental texture information at a low cost, while LiDAR can directly acquire high-precision 3D geometric information. However, using only visual sensors or LiDAR has its limitations. Visual methods are susceptible to changes in lighting and texture, and suffer from scale drift; while LiDAR provides accurate geometric information, semantic information is difficult to acquire, it is easily affected by noise, and its performance degrades in adverse weather conditions. To address the shortcomings of single sensors, vision-LiDAR fusion solutions have become mainstream, but the following problems still exist:

[0003] 1. Insufficient suppression of dynamic interference: Many systems do not have effective feature filtering mechanisms designed for dynamic environments. In scenarios with dynamic objects such as pedestrians and vehicles, dynamic feature points are easily misidentified as static reference features, resulting in significant drift in visual odometry, a sharp decrease in positioning accuracy, and insufficient robustness.

[0004] 2. Low reliability of feature matching: Traditional feature matching relies solely on descriptor similarity without incorporating semantic information for constraint. This can easily lead to numerous mismatches in areas with repetitive textures (such as long corridors or dense building clusters), further exacerbating pose estimation errors and affecting the accuracy of map construction.

[0005] 3. Loosely coupled multi-source data fusion: Most fusion schemes adopt a loosely coupled approach, only post-processing the independent positioning results of each sensor, without deeply exploring the complementarity of the two sensors. Visual point features have high accuracy but lack global structural information, while radar point clouds can provide depth information but are susceptible to noise interference. Furthermore, loose fusion cannot fully leverage the advantages of both, making it difficult to balance positioning accuracy and stability.

[0006] 4. Poor practicality of map construction: The constructed maps are mostly unordered point clouds or sparse feature maps, without incorporating semantic and structural information, which cannot provide effective structured priors for subsequent tasks, thus limiting the practical application value of the SLAM system.

[0007] Therefore, an innovative visual-radar tightly coupled positioning method is needed to effectively solve the problems mentioned above, such as dynamic interference, mismatch, loose fusion, and poor map usability, and to achieve high-precision and robust positioning and mapping in complex dynamic and structured scenarios. Summary of the Invention

[0008] The purpose of this application is to provide a method, apparatus, device, medium, and product for tightly coupled visual radar positioning of unmanned aerial vehicles (UAVs), which can improve the accuracy and robustness of positioning.

[0009] To achieve the above objectives, this application provides the following solution:

[0010] In a first aspect, this application provides a tightly coupled visual radar positioning method for unmanned aerial vehicles (UAVs), comprising:

[0011] Acquire information data; the information data includes: image data of the environment acquired by the UAV's onboard camera, and point cloud data collected by the onboard lidar;

[0012] The image data is processed using the YOLOv8-seg semantic segmentation network to obtain semantic segmentation results containing dynamic objects and structural static objects.

[0013] Based on the semantic segmentation results, feature points in the image data extracted based on visual odometry are filtered, and feature matching constraints are applied to the filtered static feature points based on the semantic consistency principle to obtain visual feature points.

[0014] The point cloud data is projected onto the image coordinate system corresponding to the image data, and semantic information is assigned to each laser point contained in the projected point cloud data according to the semantic segmentation result. Dynamic and low-confidence point clouds are filtered out according to the semantic information, and high-confidence point clouds are retained to obtain a structured semantic point cloud; wherein, the low-confidence point cloud is a point cloud with a confidence level lower than the preset confidence level.

[0015] Based on semantic-geometric consistency as a constraint, planar features are extracted from the structural semantic point cloud to obtain multiple semantic planes. These multiple semantic planes are then merged and optimized to obtain a structural semantic map of the environment.

[0016] The point-plane constraint residual is determined by transforming the visual feature points to the radar coordinate system, finding the nearest neighbor semantic plane in the structural semantic map, and calculating the vertical distance from the structural semantic point cloud and the visual feature points to the nearest neighbor semantic plane.

[0017] The observation model is determined based on the point-surface constraint residuals and semantic weight coefficients, and IESKF is used for tight-coupled filtering and optimization to obtain updated state estimation information; the updated state estimation information is used to achieve tight-coupled positioning of the UAV.

[0018] Secondly, this application provides a tightly coupled visual radar positioning device for unmanned aerial vehicles (UAVs), comprising:

[0019] An information data acquisition module is used to acquire information data; the information data includes: image data of the environment acquired by the UAV's onboard camera, and point cloud data collected by the onboard lidar.

[0020] The image processing module is used to process the image data using the YOLOv8-seg semantic segmentation network to obtain semantic segmentation results containing dynamic objects and structural static objects.

[0021] The filtering and matching module is used to filter feature points in the image data extracted based on visual odometry according to the semantic segmentation results, and to perform feature matching constraints on the filtered static feature points based on the semantic consistency principle to obtain visual feature points.

[0022] The projection processing module is used to project the point cloud data onto the image coordinate system corresponding to the image data, and assign semantic information to each laser point contained in the projected point cloud data according to the semantic segmentation result, and filter out dynamic and low-confidence point clouds according to the semantic information, retaining high-confidence point clouds to obtain structural semantic point clouds; wherein, the low-confidence point clouds are point clouds with confidence levels lower than the preset confidence level.

[0023] An extraction and optimization module is used to extract planar features from the structural semantic point cloud based on semantic-geometric consistency as a constraint, obtain multiple semantic planes, and merge and optimize the multiple semantic planes to obtain a structural semantic map of the environment.

[0024] The point-surface constraint residual determination module is used to determine the point-surface constraint residual; the point-surface constraint residual is determined by transforming the visual feature points to the radar coordinate system, finding the nearest neighbor semantic plane in the structural semantic map, and calculating the vertical distance from the structural semantic point cloud and the visual feature points to the nearest neighbor semantic plane.

[0025] The optimization estimation module is used to determine the observation model based on the point-surface constraint residuals and semantic weight coefficients, and to perform tight-coupled filtering and optimization processing using IESKF to obtain updated state estimation information; the updated state estimation information is used to achieve tight-coupled positioning of the UAV.

[0026] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described UAV visual radar tightly coupled positioning method.

[0027] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described UAV visual radar tightly coupled positioning method.

[0028] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described UAV visual radar tightly coupled positioning method.

[0029] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0030] This application provides a method, apparatus, device, medium, and product for tightly coupled visual radar localization of unmanned aerial vehicles (UAVs). By introducing semantic segmentation and point-plane tight coupling, improvements can be achieved in dynamic interference suppression, system robustness, and localization accuracy. Image data is processed using the YOLOv8-seg semantic segmentation network, and feature points are filtered based on the semantic segmentation results to eliminate dynamic interference. Simultaneously, semantic-geometric consistency is introduced as a constraint in the planar feature extraction stage, merging and optimizing multiple semantic planes to reduce the feature mismatch rate in areas with repetitive or weak textures. The observation model is determined based on the point-plane constraint residuals and semantic weight coefficients, and IESKF is used for tightly coupled filtering and optimization to obtain updated state estimation information, thereby improving localization accuracy and robustness. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 A flowchart of a tightly coupled visual radar localization method for UAVs;

[0033] Figure 2 This is a schematic diagram of a feature point filtering framework based on semantic segmentation.

[0034] Figure 3 A flowchart for point cloud projection and semantic point cloud generation;

[0035] Figure 4A flowchart for obtaining a structured semantic map from semantic point clouds;

[0036] Figure 5 A schematic diagram of radar point cloud and visual-radar measurement model;

[0037] Figure 6 This is a schematic diagram of the framework of a tightly coupled UAV vision-radar localization method based on semantic segmentation and point-area constraints.

[0038] Figure 7 A schematic diagram illustrating the operational steps of the tightly coupled visual radar positioning method for UAVs.

[0039] Figure 8 This is a structural diagram of a tightly coupled visual radar positioning device for unmanned aerial vehicles (UAVs).

[0040] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] In one exemplary embodiment, such as Figure 1 As shown, a tightly coupled visual radar localization method for UAVs is provided, including:

[0044] Step 100: Acquire information data. The information data includes: image data of the environment acquired by the UAV's onboard camera, and point cloud data collected by the onboard LiDAR.

[0045] Step 200: Use the YOLOv8-seg semantic segmentation network to process the image data to obtain semantic segmentation results containing dynamic objects and structural static objects.

[0046] Step 300: Based on the semantic segmentation results, the feature points in the image data extracted based on visual odometry are filtered, and the static feature points obtained by filtering are subject to feature matching constraints based on the semantic consistency principle to obtain visual feature points.

[0047] In one embodiment, based on the semantic segmentation results, feature points in the image data extracted based on visual odometry are filtered, and feature matching constraints are applied to the filtered static feature points based on the semantic consistency principle to obtain visual feature points, specifically including:

[0048] Based on the semantic segmentation results, feature points in the image data extracted based on visual odometry are filtered using semantic masks; the expression corresponding to the semantic mask is:

[0049] .

[0050] Based on the semantic consistency principle, feature matching constraints are applied to the selected static feature points to obtain visual feature points. The semantic consistency principle is determined using a semantic consistency function and a semantically weighted feature matching cost function. The expression for the semantic consistency function is as follows:

[0051] .

[0052] The expression for the feature matching cost function is:

[0053] .

[0054] in, For semantic masking; These are the pixel coordinates of feature points in the image data; Category labels obtained based on semantic segmentation results; Obtained based on semantic segmentation results Corresponding category tags; A collection of categories for dynamic objects; For semantic consistency functions; For the first Frame number Semantic labels for each feature point; For the first Frame number Semantic labels for each feature point; The feature matching cost function; For the first Frame number One feature point; No. Frame number One feature point; for Corresponding feature descriptors; for Corresponding feature descriptors; for and Hamming distance between them; This is a hyperparameter.

[0055] Step 400: Project the point cloud data onto the image coordinate system corresponding to the image data, and assign semantic information to each laser point in the projected point cloud data according to the semantic segmentation result. Filter out dynamic and low-confidence point clouds based on the semantic information, retaining high-confidence point clouds to obtain a structured semantic point cloud. Here, low-confidence point clouds are those with a confidence level lower than a preset confidence level; high-confidence point clouds are those with a confidence level higher than the preset confidence level. Point clouds with a confidence level equal to the preset confidence level are also considered high-confidence point clouds.

[0056] In one embodiment, point cloud data is projected onto the image coordinate system corresponding to the image data, and semantic information is assigned to each laser point contained in the projected point cloud data according to the semantic segmentation result. Dynamic and low-confidence point clouds are filtered out according to the semantic information, and high-confidence point clouds are retained to obtain a structured semantic point cloud. Specifically, this includes:

[0057] Each point in the point cloud data is transformed into the camera coordinate system using the extrinsic parameter matrix between the calibrated airborne LiDAR and the UAV's onboard camera, resulting in point cloud data in the camera coordinate system.

[0058] By using the camera intrinsic parameter matrix of the UAV's onboard camera, the point cloud data in the camera coordinate system is projected onto the image coordinate system corresponding to the image data, thus obtaining the projected point cloud data.

[0059] Based on the semantic segmentation results, semantic information is assigned to each laser point in the projected point cloud data, and point cloud filtering is performed based on the semantic probability distribution to obtain a structured semantic point cloud. The semantic probability distribution is determined using bilinear interpolation. The mathematical expression corresponding to the semantic probability distribution is:

[0060] .

[0061] in, Semantic probability distribution; The coordinates of the neighboring pixels; For interpolation weights; and All of these are circular index variables used for bilinear interpolation calculations; To The x-coordinate is rounded down; To Round down the ordinate; These are the pixel coordinates of feature points in the image data; This is the category probability distribution vector.

[0062] Step 500: Based on semantic-geometric consistency as a constraint, perform planar feature extraction on the structural semantic point cloud to obtain multiple semantic planes, and merge and optimize the multiple semantic planes to obtain a structural semantic map of the environment.

[0063] In one embodiment, semantic-geometric consistency is used as a constraint, which specifically includes a geometric proximity condition and a semantic consistency condition.

[0064] The mathematical expression corresponding to the geometric proximity condition is:

[0065] .

[0066] The mathematical expression corresponding to the semantic consistency condition is:

[0067] .

[0068] in, It is a plane normal vector; The points contained in the point set corresponding to the structural semantic point cloud; is the constant term in the plane equation; for The norm; This is a preset distance threshold; for semantic tags; Semantic labels for seed points used to generate the hypothetical plane; This is a semantic consistency function.

[0069] Step 600: Determine the point-to-surface constraint residual. The point-to-surface constraint residual is determined by transforming the visual feature points to the radar coordinate system, finding the nearest neighbor semantic plane in the structural semantic map, and calculating the vertical distance from the structural semantic point cloud and the visual feature points to the nearest neighbor semantic plane.

[0070] Step 700: Determine the observation model based on the point-area constraint residuals and semantic weight coefficients, and use IESKF for tightly coupled filtering and optimization to obtain updated state estimation information. The updated state estimation information is used to achieve tightly coupled localization of the UAV.

[0071] The mathematical expressions for the observation model include:

[0072] .

[0073] .

[0074] in, This is a visual-radar point-area observation model; semantic tags The corresponding weighting coefficients; To estimate the pose of the drone x Will The distance when projected onto the radar map plane; For the nearest neighbor semantic plane that matches in the structural semantic map Transpose of the parameters; This represents the transformation matrix from the world system to the lidar system; To estimate pose x The structural semantic point cloud after transformation to the lidar coordinate system; , All of them are the nearest neighbor semantic planes that match in the structured map. Parameters; This is a point-area observation model for lidar; To estimate the pose of the drone x Will The distance when projected onto the semantic plane; This is a radar semantic point cloud.

[0075] As an optional implementation, the observation model is determined based on the point-surface constraint residuals and semantic weight coefficients, and IESKF is used for tightly coupled filtering and optimization to obtain updated state estimation information, specifically including:

[0076] The observation model is determined based on the point-surface constraint residuals and semantic weight coefficients; a composite observation model is determined based on the observation model, and the Jacobian matrix of the composite observation model with respect to the preset error state vector is determined; the mathematical expression corresponding to the Jacobian matrix is:

[0077] .

[0078] in, It is a Jacobian matrix; It is a composite observation model; This is the preset error state vector.

[0079] Based on the Jacobian matrix, IESKF is used for tightly coupled filtering and optimization to obtain updated state estimation information.

[0080] In practical applications, such as Figure 7 As shown, the operation steps of the method mentioned in this application are as follows:

[0081] The first step is to acquire environmental image data using an airborne camera on a drone, and then process the image data using the YOLOv8-seg semantic segmentation network to obtain semantic segmentation results containing dynamic objects and structural static objects.

[0082] The second step involves filtering the feature points extracted from the visual odometry based on the semantic segmentation results. Feature points belonging to dynamic objects are removed, and the matching process for static feature points is constrained based on the semantic consistency principle. A schematic diagram of the corresponding method framework is shown below. Figure 2 .

[0083] 1. Feature points are selected by combining semantic recognition and segmentation results from the semantic segmentation results.

[0084] The recognition and segmentation results are obtained using YOLOv8-seg, and feature points are then selected by extracting the segmentation results.

[0085] .

[0086] After obtaining the semantic mask, pixel-level AND operations are performed between the semantic mask and the image to remove feature points in dynamic regions, while retaining structural feature points belonging to static regions.

[0087] 2. In the feature matching stage, assign higher weights to semantically consistent static feature point pairs.

[0088] With semantic information, we can reduce feature mismatches, such as avoiding matching feature points on a table to feature points on a computer. By matching semantically consistent feature points, we can reduce errors caused by mismatches. We define a semantically weighted feature matching cost function:

[0089] .

[0090] The higher the Cost value, the greater the likelihood that the match is incorrect. This distance measures the similarity between two points in appearance; the greater the distance, the greater the difference in appearance. It is a custom hyperparameter ( ≥0), used to control the severity of penalties for semantic inconsistency. When When =0, it degenerates into ordinary descriptor matching.

[0091] It is a semantic consistency function. If two points are semantically inconsistent (δ=0), their matching cost will be multiplied by (1+). α The cost of semantically consistent points (δ=1) remains unchanged, thus amplifying them and making them more difficult to achieve the best match. The definition is as follows:

[0092] .

[0093] Semantic labels can be obtained based on the segmentation results of the first step. .

[0094] in, These are the pixel coordinates corresponding to the i-th feature point in the t-th frame.

[0095] Figure 2 This paper presents a framework for dynamic feature filtering and visual feature point generation based on semantic segmentation. Starting with the original image input from the camera, feature points are extracted. YOLOv8-seg is then used for image recognition and semantic segmentation. Based on the segmentation results, a semantic mask is generated. Dynamic or unstable feature points are filtered out through semantic consistency checks and feature point filtering modules. Reliable feature points are then triangulated to generate visual feature points for calculating camera pose changes and motion trajectories. This method proactively eliminates dynamic interference through semantic information, improving the robustness and positioning accuracy of visual odometry in real-world complex scenes.

[0096] The third step involves projecting the point cloud data collected by the UAV's onboard LiDAR onto the image coordinate system, and assigning semantic information to each LiDAR point based on the semantic segmentation results. Dynamic and low-confidence point clouds are filtered out, while high-confidence structural semantic point clouds are retained. A schematic diagram of the corresponding method framework is shown below. Figure 3 As shown.

[0097] 1. Project the point cloud data collected by the LiDAR onto the image coordinate system to align it with the semantic segmentation results.

[0098] Let a point in the lidar point cloud be... The extrinsic parameter matrix between the lidar and the camera is obtained through calibration. (Where R is the rotation matrix and t is the translation vector) Transform it to the camera coordinate system as follows: :

[0099] .

[0100] Then, through the camera intrinsic parameter matrix K, Projected onto the image pixel coordinate system:

[0101] .

[0102] .

[0103] in, This represents the focal length of the camera along the x-axis. Let be the focal length of the camera along the y-axis. The x-coordinate of the camera principal point in the image coordinate system; Let y be the y-coordinate of the camera principal point in the image coordinate system.

[0104] 2. Based on the semantic segmentation results, assign semantic information to each laser point projected onto the image, filter out dynamic and low-confidence point clouds, and retain high-confidence structural semantic point clouds.

[0105] For point clouds that fall within the image area after projection, semantic information is obtained from the semantic segmentation results based on their pixel coordinates (u,v). Considering the density of the point cloud, projection errors caused by calibration extrinsic parameters, and semantic segmentation errors, a predefined category probability distribution vector is defined for each pixel based on the semantic segmentation results. ,in q Given the total number of categories, the precise semantic probability distribution of laser points is obtained using bilinear interpolation:

[0106] .

[0107] a is obtained from m and n The semantic probability vector of the neighboring pixels of b, The interpolation weights are determined by the sub-pixel coordinate offset and semantic category. Pixels that are closer together and semantic categories with obvious structural features have higher weights. The final weights are obtained through interpolation. This represents the semantic probability distribution vector belonging to the laser point.

[0108] Based on the obtained semantic probability distribution right Filtering is performed, where γ is the confidence threshold (γ=0.8). This process excludes points with low confidence, such as dynamic objects and unstructured regions, retaining only points with the highest semantic confidence category probability greater than 80%, such as static objects like buildings and the ground. The filtered semantic point cloud on the image is then back-projected back into the LiDAR coordinate system to form a high-confidence structured semantic point cloud C.

[0109] .

[0110] Figure 3 This paper demonstrates a semantic point cloud acquisition process based on semantic segmentation. YOLOv8-seg is used to perform semantic segmentation on the input raw image, generating segmentation results and their corresponding semantic masks. Simultaneously, the original LiDAR point cloud is projected onto the image coordinate system. Using the generated semantic masks, a bilinear interpolation algorithm is used to accurately assign semantic labels to each LiDAR point, thereby filtering out points that belong to dynamic objects or have low confidence. The final output contains only the semantic point cloud projection of high-confidence static structures. This process provides a crucial data foundation for the subsequent construction of high-quality structured semantic maps.

[0111] The fourth step is to extract planar features from the obtained structural semantic point cloud, and introduce semantic-geometric consistency as a constraint during the extraction process. Multiple extracted semantic planes are merged and optimized to construct a structural semantic map of the environment.

[0112] 1. Extract the optimal planar model from the semantic point cloud through random sampling (RANSAC) and by combining geometric and semantic constraints.

[0113] (1) Initial planar model generation.

[0114] The RANSAC algorithm is used for planar model extraction. K non-collinear points are randomly selected from the current frame's point cloud set C, and the initial planar model parameters are calculated. The planar model is defined as follows:

[0115] .

[0116] in, It is a plane normal vector. The plane intercept, Let be any point in space.

[0117] (2) Semantic-geometric consistency interior point judgment.

[0118] For each point in the point set Simultaneously, geometric proximity and semantic consistency checks are performed. A point can only be classified as an interior point of the current hypothetical plane if both conditions are met. The interior point set... The definition is as follows:

[0119] .

[0120] Geometric proximity condition: The condition is used to calculate the point. The Euclidean distance to the plane π. The distance threshold is set to a preset value. When the distance between a point and a plane is less than this threshold, the point is considered to geometrically conform to the planar model.

[0121] Semantic consistency conditions: This condition is determined by the semantic consistency function. semantic tags Is it related to the semantic labels of the seed points used to generate the hypothesis plane? The same constraint ensures that interior points are not only geometrically close but also belong to the same semantic category of objects as the seed point, effectively preventing the algorithm from incorrectly merging point clouds that are geometrically close but belong to different objects (such as a truck close to a wall).

[0122] Repeat the above sampling and interior point determination steps N times, and retain the planar model with the most interior points.

[0123] 2. The extracted planes are merged and optimized. The planes are merged and refitted by judging the angle between the normal vectors and the distance between the planes to construct a more complete and reliable structural map.

[0124] Iteratively perform the above steps on the remaining point set and the point set of each frame to extract multiple planes. Next, determine whether adjacent planes can be merged, and calculate the angle between the normal vectors of the adjacent planes. ,like and distance difference Then, the point sets in the plane that satisfy the merging conditions are merged, and the plane equation is refitted by principal component analysis to obtain the final plane.

[0125] Figure 4 This paper demonstrates the process of constructing a structured semantic map. For the obtained semantic point cloud, the RANSAC algorithm is used to extract initial planar features. Semantic-geometric consistency is introduced as a constraint to determine interior points and to merge and optimize the planes. The final output is a structured map rich in semantic information, where the environment is represented as precise planar features with category labels (such as walls and ceilings), providing a reliable environment model for subsequent accurate and robust tightly coupled localization and navigation.

[0126] The fifth step is to construct the point-plane constraint residuals between the radar point cloud and visual feature points and the radar semantic plane, combine them with semantic weights to form an observation model, and introduce it as a new tightly coupled term into the Iterative Error State Kalman Filter (IESKF) update process to optimize pose estimation by minimizing the composite residuals.

[0127] 1. Use radar point clouds and visual feature points to construct point-surface constrained residuals with planar features in structured point cloud maps.

[0128] For visual feature points in the world coordinate system after semantic filtering Estimate by current pose x Transform it to the lidar coordinate system:

[0129] .

[0130] in, The transformation matrix from the world frame to the radar frame can be obtained from the current pose and calibration parameters.

[0131] Find the semantic plane of the nearest neighbor radar point cloud for transformed visual feature points in the structural semantic map. Its plane equation is:

[0132] .

[0133] and For the nearest neighbor semantic plane that matches in the structured map The parameters.

[0134] Construct the observation model for the radar portion, and calculate the point-to-surface distance from the radar semantic point cloud to the radar point cloud semantic plane, which is the pose estimation method using the UAV. x Distance when projecting radar semantic point cloud onto a map plane:

[0135] .

[0136] To achieve tight coupling between radar and vision, an observation model is constructed that maps visual feature points to the semantic plane of the radar point cloud, i.e., through pose estimation of the UAV. x Distance when a visual point is projected onto the radar map plane:

[0137] .

[0138] Based on geometric distance calculation, semantic weight coefficients are introduced. Construct the final observation model:

[0139] .

[0140] .

[0141] semantic tags The corresponding weighting coefficients are used to enhance the contribution of static structural constraints. According to semantic probability distribution If semantic information is not used, the point-to-surface distance is directly used as the observation model. Weighting coefficients. The weighting can be dynamically adjusted based on the prior structural stability of the semantic category. Static structural objects, such as walls and the ground, are given higher weights, while dynamic or weakly structured objects are given lower weights.

[0142] The point-surface residual term is input into the IESKF framework as a semantic observation term. By minimizing this residual, the pose estimation converges towards the direction of geometric and semantic consistency between the semantic point cloud and the visual feature points on the radar semantic map, thereby achieving high-precision and drift-resistant pose calculation.

[0143] like Figure 5 As shown, Represents the environmental structural semantic plane constructed by radar, with its unit normal vector being... After transforming the visual feature points to the radar coordinate system through current pose estimation, the nearest neighbor semantic plane is found in the structural semantic map. Calculate the vertical distance from the semantic point cloud and visual feature points to the plane. and This constitutes a point-to-surface constraint residual. During the IESKF state estimation process, this residual is minimized by iteratively optimizing the pose, so that the semantic point cloud and visual feature points are precisely aligned with the radar semantic plane in terms of geometry and semantics, thereby achieving high-precision tight coupling of multi-sensor data.

[0144] 2. Tightly coupled filtering and optimization are performed using IESKF. Based on the original IMU and visual observations, radar-visual point-area constraint residuals are introduced as new observation terms.

[0145] Define the error state vector as follows:

[0146] .

[0147] in, These represent the position, velocity, and attitude errors, respectively. and These represent the zero-bias errors of the gyroscope and accelerometer, respectively. The complete state vector. Includes nominal state and error state ,satisfy:

[0148] .

[0149] Based on IMU measurements Perform state prediction Covariance propagation:

[0150] .

[0151] .

[0152] in, Let Jacobian be the state transition matrix. Here, is the noise driving matrix, and Q is the process noise covariance matrix. This step is driven solely by the IMU and is the front-end prediction stage of the system, independent of whether radar observation is incorporated.

[0153] The visual odometry framework was revised by incorporating radar-visual point-to-surface semantic constraint residuals, which were then concatenated into a composite observation vector. and its composite observation model :

[0154] .

[0155] in, These are the pixel coordinates of a feature point that the camera actually detects in the image; It is the observation model of visual odometry, which is the back-projected pixel coordinates of triangulated feature points. These are observations from the radar section, and their points... It should fall exactly on its corresponding radar semantic plane, and the distance from the point to the plane should be zero; It is a radar observation model, which estimates the current location. Below, the actual distance from the radar point cloud to the radar semantic plane. For visual-radar point-area constraints, the visual feature points It should fall exactly on its corresponding radar semantic plane, that is, the distance from the point to the plane should be zero; It is a visual-radar observation model, which calculates the estimate at the current location. Below, the actual distance from the visual feature point to the radar semantic plane.

[0156] The corresponding noise covariance matrix is:

[0157] .

[0158] in, , and These characterize the uncertainties of visual measurements, radar measurements, and visual-radar point-area constrained measurements, respectively.

[0159] Calculate the composite observation model For error state Jacobian matrix:

[0160] .

[0161] in, It is the Jacobian of the constraint term residuals with respect to the error state, which needs to be calculated using the chain rule on the transformation matrix. Differentiating the rotation and translation variables in the equation is key to achieving tight coupling of multiple sources. Let be the Jacobian matrix of the visual reprojection residual with respect to the error state; Let be the Jacobian matrix of the point-area residuals of the lidar with respect to the error state; Let be the Jacobian matrix of the visual-radar point-area residuals with respect to the error state; For visual reprojection observation models; This is a point-area observation model for lidar; This is a visual-radar point-area observation model.

[0162] The standard iterative update loop for IESKF using composite observations is as follows:

[0163] Residual calculation:

[0164] .

[0165] in, For residuals; This is the observation model.

[0166] Calculate Kalman gain :

[0167] .

[0168] Update error status:

[0169] .

[0170] Updated state estimate:

[0171] .

[0172] Repeat the above steps until convergence. ).

[0173] After the iteration converges, update the final state and covariance:

[0174] .

[0175] .

[0176] in, The covariance matrix of the observation noise; This is the optimal estimate of the system state vector; For update operator; Let be the covariance matrix of the state estimation error at time k+1; It is the identity matrix; The Jacobian matrix of the observation model for the state vector; This is the state error covariance matrix before the iterative update.

[0177] The IESKF framework introduces radar point-area constraint residuals as tightly coupled observation terms, deeply integrating the accuracy of visual point features, the dynamic responsiveness of the IMU, and the global structurality of radar surface features. It uses semantic information weighting to ensure the reliability of constraints, thereby achieving high-precision pose estimation while effectively suppressing the cumulative drift error, especially in structured scenarios such as long corridors and open squares.

[0178] Figure 6This paper presents a framework for a tightly coupled visual-radar localization method for UAVs based on semantic segmentation and point-area constraints. The core idea is to deeply fuse multi-sensor data around the IESKF framework and improve the system's accuracy and robustness through semantic information. The process begins with multi-sensor data input, introducing a semantic segmentation module to identify dynamic and static structural objects in images and point clouds. The system then uses this semantic information to perform two key operations: first, filtering visual feature points to actively remove unreliable features and avoid localization drift; second, extracting semantic point clouds from radar point clouds and forming planar structures with semantic labels to construct a structured semantic map. Finally, all this information is combined to construct a residual term, which is input into the IESKF framework for optimal state estimation, thereby outputting high-precision, drift-resistant pose results in complex dynamic environments.

[0179] By introducing semantic segmentation and point-to-surface tight coupling constraints, improvements were achieved in dynamic interference suppression, system robustness, and positioning accuracy. At the data front-end processing level, the YOLOv8-seg semantic segmentation network accurately identifies object categories and generates semantic masks to filter out unreliable visual feature points, eliminating the interference of dynamic disturbances on visual odometry. Simultaneously, a semantic consistency weighted cost function was introduced in the feature matching stage, assigning higher weights to semantically consistent feature pairs, reducing the feature mismatch rate in areas with repetitive or weak textures.

[0180] Deep, tight coupling was achieved at the multi-source data fusion and pose estimation optimization levels. By aligning and filtering radar point clouds with semantic information, high-confidence structural semantic point clouds were preserved, providing high-quality input for fusion. Point-area constraint residuals between radar point clouds, visual feature points, and the radar semantic plane were constructed and used as new observation terms, inputting them along with IMU and visual observations into an iterative error state Kalman filter framework for optimization. This fully leverages the accuracy of visual point features, the dynamism of the IMU, and the global structural advantages of radar surface features, effectively utilizing the structural constraints of the environment. In challenging scenarios such as open squares lacking texture features or with ambiguous scales, long-distance corridors, the long-term cumulative drift error was reduced, improving pose estimation accuracy and system stability.

[0181] Through semantic plane extraction and optimization processes, a structured semantic map with practical application value was generated. This map consists of optimized planes with semantic labels, rather than a disordered point cloud. This structured semantic map provides directly usable prior information for subsequent advanced tasks of UAVs, such as path planning, obstacle avoidance, and scene understanding, greatly expanding the practical application potential and system intelligence level of SLAM systems in fields such as autonomous navigation and intelligent inspection.

[0182] In one exemplary embodiment, such as Figure 8As shown, a tightly coupled visual radar positioning device for unmanned aerial vehicles (UAVs) is provided, comprising:

[0183] The information data acquisition module is used to acquire information data. This information data includes: environmental image data acquired by the UAV's onboard camera, and point cloud data collected by the onboard LiDAR.

[0184] The image processing module is used to process image data using the YOLOv8-seg semantic segmentation network to obtain semantic segmentation results containing dynamic objects and structural static objects.

[0185] The filtering and matching module is used to filter feature points in image data extracted based on visual odometry according to semantic segmentation results, and to perform feature matching constraints on the filtered static feature points based on the semantic consistency principle to obtain visual feature points.

[0186] The projection processing module projects point cloud data onto the image coordinate system corresponding to the image data. Based on the semantic segmentation results, it assigns semantic information to each laser point in the projected point cloud data. Then, it filters out dynamic and low-confidence point clouds based on the semantic information, retaining high-confidence point clouds to obtain a structured semantic point cloud. The low-confidence point clouds are those with a confidence level lower than a preset confidence level.

[0187] The extraction and optimization module is used to extract planar features from the structural semantic point cloud based on semantic-geometric consistency as a constraint, obtain multiple semantic planes, and merge and optimize the multiple semantic planes to obtain a structural semantic map of the environment.

[0188] The point-area constraint residual determination module is used to determine the point-area constraint residuals. The point-area constraint residuals are determined by transforming the visual feature points to the radar coordinate system, finding the nearest neighbor semantic plane in the structured semantic map, and calculating the perpendicular distance from the structured semantic point cloud and the visual feature points to the nearest neighbor semantic plane.

[0189] The optimization estimation module determines the observation model based on point-area constraint residuals and semantic weight coefficients, and uses IESKF for tightly coupled filtering and optimization to obtain updated state estimation information. This updated state estimation information is then used to achieve tightly coupled localization of the UAV.

[0190] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 9As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores UAV visual radar tightly coupled positioning data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the UAV visual radar tightly coupled positioning method.

[0191] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0192] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0193] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0194] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0195] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0196] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0197] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logic devices, etc., and are not limited to these.

[0198] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0199] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A tightly coupled visual radar positioning method for unmanned aerial vehicles (UAVs), characterized in that, include: Acquire information and data; The information data includes: image data of the environment acquired by the UAV's onboard camera, and point cloud data collected by the onboard lidar. The image data is processed using the YOLOv8-seg semantic segmentation network to obtain semantic segmentation results containing dynamic objects and structural static objects. Based on the semantic segmentation results, feature points in the image data extracted based on visual odometry are filtered, and feature matching constraints are applied to the filtered static feature points based on the semantic consistency principle to obtain visual feature points. The point cloud data is projected onto the image coordinate system corresponding to the image data, and semantic information is assigned to each laser point contained in the projected point cloud data according to the semantic segmentation result. Dynamic and low-confidence point clouds are filtered out according to the semantic information, and high-confidence point clouds are retained to obtain a structured semantic point cloud; wherein, the low-confidence point cloud is a point cloud with a confidence level lower than the preset confidence level. Based on semantic-geometric consistency as a constraint, planar features are extracted from the structural semantic point cloud to obtain multiple semantic planes. These multiple semantic planes are then merged and optimized to obtain a structural semantic map of the environment. The point-plane constraint residual is determined by transforming the visual feature points to the radar coordinate system, finding the nearest neighbor semantic plane in the structural semantic map, and calculating the vertical distance from the structural semantic point cloud and the visual feature points to the nearest neighbor semantic plane. The observation model is determined based on the point-surface constraint residuals and semantic weight coefficients, and IESKF is used for tight-coupled filtering and optimization to obtain updated state estimation information; the updated state estimation information is used to achieve tight-coupled positioning of the UAV.

2. The UAV visual radar tightly coupled positioning method according to claim 1, characterized in that, Based on the semantic segmentation results, feature points in the image data extracted based on visual odometry are filtered, and feature matching constraints are applied to the filtered static feature points based on the semantic consistency principle to obtain visual feature points, specifically including: Based on the semantic segmentation results, feature points in the image data extracted based on visual odometry are filtered using a semantic mask; the expression corresponding to the semantic mask is: ; Based on the semantic consistency principle, feature matching constraints are applied to the selected static feature points to obtain visual feature points; the semantic consistency principle is determined using a semantic consistency function and a semantically weighted feature matching cost function; the expression of the semantic consistency function is: ; The expression for the feature matching cost function is: ; in, For semantic masking; These are the pixel coordinates of feature points in the image data; Category labels obtained based on semantic segmentation results; Obtained based on semantic segmentation results Corresponding category tags; A collection of categories for dynamic objects; For semantic consistency functions; For the first Frame number Semantic labels for each feature point; For the first Frame number Semantic labels for each feature point; The feature matching cost function; For the first Frame number One feature point; No. Frame number One feature point; for Corresponding feature descriptors; for Corresponding feature descriptors; for and Hamming distance between them; This is a hyperparameter.

3. The UAV visual radar tightly coupled positioning method according to claim 1, characterized in that, The point cloud data is projected onto the image coordinate system corresponding to the image data, and semantic information is assigned to each laser point contained in the projected point cloud data according to the semantic segmentation result. Dynamic and low-confidence point clouds are filtered out according to the semantic information, and high-confidence point clouds are retained to obtain a structured semantic point cloud, specifically including: Each point in the point cloud data is transformed into the camera coordinate system by the extrinsic parameter matrix between the airborne lidar and the UAV's airborne camera obtained through calibration, to obtain point cloud data in the camera coordinate system. By using the camera intrinsic parameter matrix of the UAV's onboard camera, the point cloud data in the camera coordinate system is projected onto the image coordinate system corresponding to the image data to obtain the projected point cloud data. Based on the semantic segmentation results, semantic information is assigned to each laser point in the projected point cloud data, and point cloud filtering is performed based on the semantic probability distribution to obtain a structured semantic point cloud; the semantic probability distribution is determined using bilinear interpolation; the mathematical expression corresponding to the semantic probability distribution is: ; in, Semantic probability distribution; The coordinates of the neighboring pixels; For interpolation weights; and All of these are circular index variables used for bilinear interpolation calculations; To The x-coordinate is rounded down; To Round down the ordinate; These are the pixel coordinates of feature points in the image data; This is the category probability distribution vector.

4. The UAV visual radar tightly coupled positioning method according to claim 1, characterized in that, Based on semantic-geometric consistency as a constraint, the constraint specifically includes: geometric proximity condition and semantic consistency condition; The mathematical expression corresponding to the geometric proximity condition is: ; The mathematical expression corresponding to the semantic consistency condition is: ; in, It is a plane normal vector; The points contained in the point set corresponding to the structural semantic point cloud; is the constant term in the plane equation; for The norm; This is a preset distance threshold; for semantic tags; Semantic labels for seed points used to generate the hypothetical plane; This is a semantic consistency function.

5. The UAV visual radar tightly coupled positioning method according to claim 1, characterized in that, The mathematical expression of the observation model includes: ; ; in, This is a visual-radar point-area observation model; semantic tags The corresponding weighting coefficients; To estimate the pose of the drone x Will The distance when projected onto the radar map plane; For the nearest neighbor semantic plane that matches in the structural semantic map Transpose of the parameters; This represents the transformation matrix from the world system to the lidar system; To estimate pose x The structural semantic point cloud after transformation to the lidar coordinate system; , All of them are the nearest neighbor semantic planes that match in the structured map. Parameters; This is a point-area observation model for lidar; To estimate the pose of the drone x Will The distance when projected onto the semantic plane; This is a radar semantic point cloud.

6. The UAV visual radar tightly coupled positioning method according to claim 1, characterized in that, The observation model is determined based on the point-surface constraint residuals and semantic weight coefficients, and IESKF is used for tightly coupled filtering and optimization to obtain updated state estimation information, specifically including: The observation model is determined based on the point-surface constraint residuals and semantic weight coefficients. Based on the observation model, a composite observation model is determined, and the Jacobian matrix of the composite observation model for the preset error state vector is determined; the mathematical expression corresponding to the Jacobian matrix is: ; in, It is a Jacobian matrix; It is a composite observation model; This is a preset error state vector; Based on the Jacobian matrix, IESKF is used for tightly coupled filtering and optimization to obtain updated state estimation information.

7. A tightly coupled positioning device for UAV visual radar, characterized in that, include: Information data acquisition module, used to acquire information data; The information data includes: image data of the environment acquired by the UAV's onboard camera, and point cloud data collected by the onboard lidar. The image processing module is used to process the image data using the YOLOv8-seg semantic segmentation network to obtain semantic segmentation results containing dynamic objects and structural static objects. The filtering and matching module is used to filter feature points in the image data extracted based on visual odometry according to the semantic segmentation results, and to perform feature matching constraints on the filtered static feature points based on the semantic consistency principle to obtain visual feature points. The projection processing module is used to project the point cloud data onto the image coordinate system corresponding to the image data, and assign semantic information to each laser point contained in the projected point cloud data according to the semantic segmentation result, and filter out dynamic and low-confidence point clouds according to the semantic information, retaining high-confidence point clouds to obtain structural semantic point clouds; wherein, the low-confidence point clouds are point clouds with confidence levels lower than the preset confidence level. An extraction and optimization module is used to extract planar features from the structural semantic point cloud based on semantic-geometric consistency as a constraint, obtain multiple semantic planes, and merge and optimize the multiple semantic planes to obtain a structural semantic map of the environment. The point-surface constraint residual determination module is used to determine the point-surface constraint residual; the point-surface constraint residual is determined by transforming the visual feature points to the radar coordinate system, finding the nearest neighbor semantic plane in the structural semantic map, and calculating the vertical distance from the structural semantic point cloud and the visual feature points to the nearest neighbor semantic plane. The optimization estimation module is used to determine the observation model based on the point-surface constraint residuals and semantic weight coefficients, and to perform tight-coupled filtering and optimization processing using IESKF to obtain updated state estimation information; the updated state estimation information is used to achieve tight-coupled positioning of the UAV.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the UAV visual radar tightly coupled positioning method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the UAV visual radar tightly coupled positioning method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the UAV visual radar tightly coupled positioning method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Joint optimization dynamic SLAM method based on semantics and geometry

    CN112308921A

  • Dynamic environment semantic SLAM method based on monocular vision and LiDAR fusion

    CN119863578A