Semi-dynamic obstacle identification method and device for unmanned aerial platform

By processing high-resolution color images and point cloud data of the unmanned aerial platform and extracting the multi-dimensional features of obstacles, the problem of the unmanned aerial platform identifying and navigating semi-dynamic obstacles in complex environments is solved, and accurate obstacle avoidance decisions and real-time path planning are achieved.

CN120673414APending Publication Date: 2025-09-19BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510773722.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing unmanned aerial platforms have difficulty effectively handling semi-dynamic obstacles, especially when traveling inside high-rise buildings and confined spaces. They are unable to achieve efficient identification and interactive navigation of semi-dynamic obstacles.

Method used

By acquiring high-resolution color images and corresponding ordered point cloud data, open-set target detection and visual segmentation are performed to extract the multi-dimensional features of obstacles, including the oriented bounding boxes and orientation vectors of static, dynamic, and semi-dynamic obstacles. Path planning is performed by combining the RRT, A*, and Minimal Snap algorithms to determine the passability of obstacles.

Benefits of technology

It achieves accurate recognition of semi-dynamic obstacles and interactive navigation capabilities, provides accurate obstacle avoidance decisions for unmanned aerial platforms, and meets the needs of real-time and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673414A_ABST
    Figure CN120673414A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-dynamic obstacle recognition method and device for an unmanned flight platform, and relates to the technical field of obstacle recognition and trajectory planning, and the method comprises the steps: carrying out the open set target detection of a color image, obtaining a bounding box of an obstacle, and carrying out the visual segmentation of the color image and the bounding box, and obtaining a segmentation mask of the obstacle; obtaining a point cloud cluster set corresponding to the obstacle based on the ordered point cloud data and the segmentation mask; when the obstacles comprise a non-semi-dynamic obstacle and a semi-dynamic obstacle, calculating a directional bounding box for the point cloud cluster set of the non-semi-dynamic obstacle, and fitting an orientation vector for the point cloud cluster set of the semi-dynamic obstacle; obtaining semi-dynamic obstacle state information based on the orientation vector of the semi-dynamic obstacle and the orientation vector of the wall body; and according to the semi-dynamic obstacle state information and the orientation bounding box of the non-semi-dynamic obstacle, obtaining a planning track. According to the invention, efficient identification and interactive passing of the semi-dynamic obstacle can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of obstacle identification and trajectory planning, and in particular to a semi-dynamic obstacle identification method and device for an unmanned aerial platform. Background Art

[0002] Traditional unmanned aerial vehicles (UAVs) are primarily designed for operations in open spaces. Their environmental perception systems are centered around conventional obstacle avoidance, utilizing sensors like lidar and cameras to achieve obstacle avoidance and path planning. With the growing demand for complex indoor operations in areas like emergency rescue and facility inspection, the application of UAVs is expanding into confined environments like the interiors of high-rise buildings and confined spaces. These indoor environments are often plagued by unique obstacles like doors, windows, nets, and curtains—known as semi-dynamic obstacles—which pose new challenges that existing obstacle avoidance systems struggle to address.

[0003] Different from the static obstacles or moving dynamic obstacles in traditional obstacle avoidance scenarios, semi-dynamic obstacles remain stationary under normal conditions, but can be displaced by specific external forces. In scenarios such as emergency rescue and facility inspections, unmanned aerial platforms often face the need to navigate through semi-dynamic obstacles (such as passing through half-closed doors and windows to enter confined spaces during post-disaster search and rescue). This poses a new challenge to existing obstacle avoidance systems - not only do they need to achieve conventional obstacle avoidance, but they also need to be able to judge the passability of obstacles and interactively navigate through them. However, the obstacle avoidance technology of existing unmanned aerial platforms is mainly designed for static obstacles or dynamic obstacles, and it is difficult to effectively handle the complex characteristics of semi-dynamic obstacles, and a breakthrough solution is urgently needed. Summary of the Invention

[0004] The purpose of this application is to provide a semi-dynamic obstacle recognition method and device for an unmanned aerial platform, which can achieve efficient recognition and interactive navigation of semi-dynamic obstacles.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a semi-dynamic obstacle recognition method for an unmanned aerial platform, comprising:

[0007] Obtain high-resolution color images and corresponding ordered point cloud data;

[0008] performing open set object detection on the color image to obtain a bounding box of the obstacle, and performing visual segmentation on the color image and the bounding box to obtain a segmentation mask of the obstacle; the obstacle includes a non-semi-dynamic obstacle and a semi-dynamic obstacle; the non-semi-dynamic obstacle is a static obstacle or a dynamic obstacle, or the non-semi-dynamic obstacle includes a static obstacle and a dynamic obstacle;

[0009] Obtaining point cloud clusters corresponding to obstacles based on the ordered point cloud data and the segmentation mask;

[0010] When the obstacles include both non-semi-dynamic and semi-dynamic obstacles, an oriented bounding box is calculated for the point cloud clusters of the non-semi-dynamic obstacles, and a heading vector is fitted for the point cloud clusters of the semi-dynamic obstacles. Then, based on the heading vectors of the semi-dynamic obstacles and the obtained heading vectors of the walls, the semi-dynamic obstacle state information is obtained. Finally, based on the semi-dynamic obstacle state information and the oriented bounding boxes of the non-semi-dynamic obstacles, the passability of all obstacles is determined to obtain the planned trajectory of the unmanned aerial platform.

[0011] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned semi-dynamic obstacle recognition method for an unmanned aerial platform.

[0012] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned semi-dynamic obstacle recognition method for an unmanned aerial platform.

[0013] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned semi-dynamic obstacle recognition method for an unmanned aerial platform.

[0014] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0015] This application provides a method and device for identifying semi-dynamic obstacles for unmanned aerial platforms. By performing open-set target detection and visual segmentation on high-resolution color images, accurate bounding box positioning and pixel-level mask segmentation of obstacles are achieved, thereby effectively distinguishing static, dynamic, and semi-dynamic obstacles. Moreover, by extracting different multi-dimensional features (i.e., semi-dynamic obstacle state information, directional bounding boxes of static and dynamic obstacles) based on the dynamic characteristics of the obstacles (static, dynamic, semi-dynamic), it is possible to accurately judge the passability of obstacles based on the different multi-dimensional features of the obstacles, and also achieve the interactive traversal capability of semi-dynamic obstacles, thereby providing accurate obstacle avoidance decisions for unmanned aerial platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is a flow chart of a semi-dynamic obstacle recognition method for an unmanned aerial platform in Example 1 of the present application;

[0018] Figure 2 This is a structural diagram of the unmanned aerial platform equipment in Example 2 of this application;

[0019] Figure 3 This is a diagram of the software package architecture provided in Example 2 of this application;

[0020] Figure 4 This is a flow chart of the real-time detection and planning method for multi-dimensional features of obstacles provided in Example 2 of this application;

[0021] Figure 5 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] Currently, unmanned aerial platforms, such as small and micro drones, can avoid static and dynamic obstacles through existing obstacle avoidance technology. However, when faced with scenes blocked by semi-dynamic obstacles such as doors, windows, nets and curtains in buildings, they can only treat semi-dynamic obstacles as static obstacles and avoid them, while ignoring their interactive features, reflecting a single obstacle avoidance strategy. Although some obstacle detection methods for complex environments have been proposed, their technical routes have the following significant limitations: (1) The interactive feature detection capability is insufficient. The identification of semi-dynamic obstacles requires a comprehensive analysis of their structural variability and state dependence: first, the deformable features of the obstacles (such as the rotation of door leaves around the axis and the swinging movement of curtains) must be analyzed rather than simply modeled as fixed geometric bodies; second, the interactive state (such as the opening and closing angles of doors and windows, the transparency of net curtains) must be judged in real time, which directly affects the feasibility of passage; third, the motion constraint mechanism (such as the hinge position and the direction of the slide rail) must be identified to predict its kinematic behavior; at the same time, the material properties (such as the transparency of glass doors and the flexibility of gauze) directly affect the sensor perception characteristics and collision risk level; finally, the semantic category of semi-dynamic obstacles is strongly coupled with their physical state (such as an obstacle when the door is closed and an open space when it is opened), requiring the system to realize the dynamic fusion of geometric characteristics and semantic understanding. Existing methods simplify obstacles into static grids or dynamic trajectories and cannot capture such complex coupling relationships. (2) Limited versatility: Existing detection methods often use closed-set self-supervised models, and the detection categories are limited by the training dataset, making it difficult to adapt to the diversity of real-world scenarios. (3) Real-time defects: Existing methods often cannot meet the real-time requirements of UAV movement and operation due to the high computational overhead caused by the complex architecture when processing complex scenarios. Therefore, this application provides a semi-dynamic obstacle recognition method for UAVs to address the above-mentioned shortcomings.

[0024] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0025] Example 1

[0026] In an exemplary embodiment, Figure 1 As shown, a semi-dynamic obstacle recognition method for an unmanned aerial platform is provided. The method is executed by a computer device integrated in a terminal, wherein the terminal can be, but is not limited to, various unmanned aerial platforms. In the embodiment of the present application, the method is described by applying the computer device as an example, and includes the following steps 101 to 104. Among them:

[0027] Step 101: Obtain a high-resolution color image and corresponding ordered point cloud data of the current scene.

[0028] In step 102, open set object detection is performed on the color image to obtain a bounding box of the obstacle, and visual segmentation is performed on the color image and the bounding box to obtain a segmentation mask of the obstacle. Obstacles include non-semi-dynamic obstacles and semi-dynamic obstacles. Non-semi-dynamic obstacles are static obstacles or dynamic obstacles, or non-semi-dynamic obstacles include static obstacles and dynamic obstacles.

[0029] Step 103: Obtain point cloud clusters corresponding to obstacles based on the ordered point cloud data and the segmentation mask.

[0030] In step 104, when the obstacles include both non-semi-dynamic and semi-dynamic obstacles, an oriented bounding box is calculated for the point cloud clusters of the non-semi-dynamic obstacles, and an orientation vector is fitted for the point cloud clusters of the semi-dynamic obstacles. Then, based on the orientation vectors of the semi-dynamic obstacles and the obtained orientation vectors of the walls, the semi-dynamic obstacle state information is obtained. Finally, based on the semi-dynamic obstacle state information and the oriented bounding boxes of the non-semi-dynamic obstacles, the passability of all obstacles is determined to obtain the planned trajectory of the unmanned aerial platform.

[0031] Specifically, step 104 includes three feasible technical solutions: The first solution: When the obstacles include both static and semi-dynamic obstacles, an oriented bounding box is calculated for the point cloud clusters of the static obstacles, and an orientation vector is fitted for the point cloud clusters of the semi-dynamic obstacles. Then, based on the orientation vectors of the semi-dynamic obstacles and the obtained orientation vectors of the walls, the semi-dynamic obstacle state information is obtained. Finally, based on the semi-dynamic obstacle state information and the oriented bounding boxes of the static obstacles, a passability determination is performed for all obstacles to obtain the planned trajectory of the unmanned aerial vehicle platform. The second solution: When the obstacles include both dynamic and semi-dynamic obstacles, an oriented bounding box is calculated for the point cloud clusters of the dynamic obstacles, and an orientation vector is fitted for the point cloud clusters of the semi-dynamic obstacles. Then, based on the orientation vectors of the semi-dynamic obstacles and the obtained orientation vectors of the walls, the semi-dynamic obstacle state information is obtained. Finally, based on the semi-dynamic obstacle state information and the oriented bounding boxes of the dynamic obstacles, a passability determination is performed for all obstacles to obtain the planned trajectory of the unmanned aerial vehicle platform. The third solution: When the obstacles include static, dynamic, and semi-dynamic obstacles, oriented bounding boxes are calculated for the point cloud clusters of the static and dynamic obstacles, and orientation vectors are fitted for the point cloud clusters of the semi-dynamic obstacles. The semi-dynamic obstacle status information is then obtained based on the orientation vectors of the semi-dynamic obstacles and the obtained orientation vectors of the walls. Finally, the passability of all obstacles is determined based on the semi-dynamic obstacle status information and the oriented bounding boxes of the static and dynamic obstacles to obtain the planned trajectory of the unmanned aerial platform.

[0032] By implementing steps 101 to 104 above, open-set object detection and visual segmentation are performed on high-resolution color images, achieving precise bounding box positioning and pixel-level mask segmentation of obstacles, effectively distinguishing between static, dynamic, and semi-dynamic obstacles. Furthermore, by extracting different multidimensional features (i.e., semi-dynamic obstacle state information, and oriented bounding boxes for static and dynamic obstacles) based on the dynamic characteristics of the obstacles (static, dynamic, and semi-dynamic), the passability of the obstacles can be accurately determined based on their different multidimensional features. This also enables interactive navigation through semi-dynamic obstacles, providing precise obstacle avoidance decisions for unmanned aerial vehicles.

[0033] Furthermore, in step 102, open set object detection is performed using an optimization model based on OWL-ViT. Open set object detection is performed on the color image to obtain the bounding box of the obstacle, specifically including:

[0034] a. Divide the color image into image blocks, embed the image blocks into vector sequences using linear mapping, and encode the vector sequences using a multi-layer visual transformer to obtain a global feature map.

[0035] b. Use the text encoder to transform the obtained obstacle category to obtain a text feature vector.

[0036] c. Use the cross-attention mechanism to perform cross-modal feature alignment on the global feature map and text feature vector to obtain aligned multimodal features.

[0037] d. Generate candidate bounding boxes based on multimodal features, use non-maximum suppression to filter candidate bounding boxes, and output the bounding box of the obstacle.

[0038] Furthermore, in step 102, the visual segmentation is performed using a model based on SAM project distillation; the model based on SAM project distillation includes an encoder and a decoder; the encoder includes a visual backbone network and a hint encoding module; visual segmentation is performed on the color image and the bounding box to obtain a segmentation mask of the obstacle, specifically including:

[0039] a. Input the color image into the visual backbone network to obtain a multi-scale feature map.

[0040] b. Use the hint encoding module to associate the bounding box and the multi-scale feature map to obtain the association result.

[0041] c. Use the decoder to upsample the association results to obtain a multi-channel probability map, and convert the multi-channel probability map into a segmentation mask of the obstacle.

[0042] Furthermore, in step 103 , point cloud clusters corresponding to obstacles are obtained based on the ordered point cloud data and the segmentation mask, which specifically includes: processing the ordered point cloud data using a point-by-point traversal method according to the segmentation mask to obtain point cloud clusters corresponding to the obstacles.

[0043] Furthermore, in step 104, the oriented bounding box is calculated for the point cloud cluster of the non-semi-dynamic obstacle, specifically including:

[0044] a. Perform point cloud preprocessing on the point cloud cluster of the non-semi-dynamic obstacle to obtain a first point cloud preprocessing result; the point cloud preprocessing includes filtering denoising and clustering segmentation.

[0045] b. Perform first-order statistics calculation on the first point cloud preprocessing result to obtain the centroid coordinates, and perform second-order statistics calculation on the centroid coordinates to obtain the covariance matrix.

[0046] c. Perform eigendecomposition on the covariance matrix to obtain the main direction vector, and project the first point cloud preprocessing result onto the main direction vector to obtain the oriented bounding box.

[0047] Furthermore, in step 104, the orientation vector is fitted to the point cloud cluster of the semi-dynamic obstacle, which specifically includes:

[0048] a. Perform accelerated point cloud preprocessing on the point cloud cluster of the semi-dynamic obstacle to obtain a second point cloud preprocessing result; the accelerated point cloud preprocessing includes downsampling and outlier removal.

[0049] b. Use the RANSAC algorithm to perform plane fitting and normal extraction on the second point cloud preprocessing results; obtain the orientation vector of the semi-dynamic obstacle.

[0050] Furthermore, in step 104, the semi-dynamic obstacle includes a rigid semi-dynamic obstacle and / or a flexible semi-dynamic obstacle; the semi-dynamic obstacle state information includes the opening and closing angle of the rigid semi-dynamic obstacle and / or the connection state of the flexible semi-dynamic obstacle. The semi-dynamic obstacle state information is obtained based on the orientation vector of the semi-dynamic obstacle and the obtained orientation vector of the wall, specifically including:

[0051] a. Remove the obstacle point cloud data from the ordered point cloud data to obtain the wall point cloud data.

[0052] b. Use the RANSAC algorithm to perform plane fitting and normal extraction on the wall point cloud data to obtain the wall's orientation vector.

[0053] c. If the semi-dynamic obstacle is rigid, use the quaternion method to calculate the orientation vector of the rigid semi-dynamic obstacle and the orientation vector of the corresponding wall to obtain a quaternion. Then perform Euler angle decomposition and opening and closing angle determination on the quaternion to obtain the opening and closing angle of the rigid semi-dynamic obstacle.

[0054] d. If the semi-dynamic obstacle is a flexible obstacle, calculate the average normal vector of the obstacle's orientation vector. Based on the obstacle's orientation vector and the average normal vector, calculate the dispersion index of the obstacle to determine the connection status of the obstacle.

[0055] e. When the semi-dynamic obstacle includes both rigid and flexible semi-dynamic obstacles, the quaternion method is used to calculate the orientation vector of the rigid semi-dynamic obstacle and the orientation vector of the corresponding wall to obtain a quaternion. The quaternion is then decomposed into Euler angles and the opening and closing angles are determined to obtain the opening and closing angles of the rigid semi-dynamic obstacle. The average normal vector of the orientation vector of the flexible semi-dynamic obstacle is calculated. Based on the orientation vector and average normal vector of the flexible semi-dynamic obstacle, the discreteness index of the flexible semi-dynamic obstacle is calculated to determine the connection status of the flexible semi-dynamic obstacle.

[0056] Furthermore, in step 104, the passability of all obstacles is determined based on the semi-dynamic obstacle state information and the oriented bounding boxes of the non-semi-dynamic obstacles to obtain the planned trajectory of the unmanned aerial platform, which specifically includes:

[0057] a. Based on the oriented bounding box of the non-semi-dynamic obstacle, the first trajectory is generated using the RRT algorithm or the A* algorithm.

[0058] b. Based on the semi-dynamic obstacle state information, a second trajectory is generated using the trajectory generation algorithm of the interaction point and preview path.

[0059] c. Use Minimal Snap trajectory optimization to optimize the second trajectory to obtain the third trajectory.

[0060] d. Based on the first trajectory and the third trajectory, determine the planned trajectory.

[0061] Example 2

[0062] The following takes the third solution in step 104 as an example to illustrate the semi-dynamic obstacle recognition method for the unmanned aerial platform in detail.

[0063] When an unmanned aerial vehicle (UAV) performs search, reconnaissance, and patrol missions in complex indoor environments, it implements a hierarchical response decision-making process upon encountering an obstacle. First, in the far-field phase (detection distance > 5m), the visual sensor performs a rough obstacle screening, quickly extracting basic information such as obstacle type and target category. Then, after the UAV approaches the obstacle to a moderate distance, it performs obstacle recognition and trajectory planning. This obstacle recognition and planning method includes the following steps:

[0064] Step 1: The depth camera obtains the color image and ordered point cloud data of the current scene, and performs open set image segmentation based on the current frame color image data.

[0065] Use a depth camera that can reconstruct 3D point clouds from depth images to obtain high-resolution color images and corresponding color ordered point cloud data. o (i.e. name) as a text prompt, together with the color image, is input into the open set target detection model optimized based on the OWL-ViT (Vision Transformer for Open-World Localization) project, and the bounding box of the corresponding obstacle is obtained. OWL-ViT realizes open set target detection by combining visual transformers and multimodal semantic embedding. The specific process of obstacle bounding box detection can be divided into the following core steps. The first is image input and feature extraction: the input color image is divided into regular image blocks, each image block is embedded into a vector sequence through linear mapping, and position encoding is added to retain spatial information; this vector sequence is passed through a multi-layer visual transformer (ViT) g v (·) is used for encoding. Each layer captures the global contextual relationship through a multi-head self-attention mechanism, gradually abstracting the high-level semantic features of the image to form a global feature map. The second is text-driven semantic embedding: In the open set scenario, the obstacle category l is input. o Through the text encoder g t (·) is converted into a text feature vector The corresponding d-dimensional vector space, d is the dimension of text embedding, and different obstacle categories correspond to different semantic embeddings. This process enables the model to dynamically match text semantics with image content, supporting plug-and-play detection of any category without retraining. Next is cross-modal feature alignment: a cross-attention mechanism is introduced between image features (i.e., global feature maps) and text embeddings (text feature vectors). That is, text embedding is used as the query (Q), image features as the key (K) and value (V), d is the dimension of the aforementioned image feature embedding and text embedding, that is, the length of the matrix corresponding to the query and key, T is the transpose symbol, and the correlation between the two is calculated through attention weights; this mechanism focuses on the areas in the image features that highly match the text semantics, thereby establishing a mapping relationship between the text description and the visual area in the feature space. Finally, bounding box prediction and confidence scoring are performed: Based on the aligned multimodal features, the model generates candidate bounding boxes from the image area, and screens the overlapping bounding boxes through non-maximum suppression (NMS), retaining the prediction boxes with the highest confidence s and independent of each other. Among them, s i is the confidence level of the predicted box, x i ,y i is the coordinate of the upper left point of the prediction box, w i ,h i To predict the bounding box width and height, τ is the confidence threshold. This step ensures that the final bounding box output is accurate and free of redundancy, resolving the issue of multiple box interference in adjacent obstacle detection. During deployment, NVIDIA's open-source NanoOWL project was used. The image encoder of the OWL-ViT model was converted to half-floating-point precision (FP16) to improve detection efficiency. The text encoder and projection models used the original models and were exported using the ONNX neural network model universal format. TensorRT was used to build the model engine to accelerate on-device model inference.

[0066] Then, the resulting bounding box As visual input, it is input into the visual segmentation model based on SAM project distillation together with the color image, and the image segmentation mask M corresponding to the obstacle is obtained. t (u, v), the specific process of the SAM model corresponding to the output image segmentation mask can be divided into the following steps. First, input the coordinate information of a single bounding box B rectangular area b = (x min ,y min ,x max ,y max ), the information is first mapped into a high-dimensional vector through position encoding PE(·) is the position encoding mapping conversion relationship, and is combined with the image encoder from the image Extracted image features F∈R H×W×C Specifically, the encoder of SAM includes a visual backbone network (Backbone) and a prompt encoding module: the image passes through the visual backbone network and other structures to generate a multi-scale feature map F = Backbone (I) ∈ R H×W×C , and the coordinates of the bounding box are converted into vectors through the embedding layer and then injected into a lightweight hint encoding module. At this time, the model uses the cross attention mechanism The potential object boundaries that overlap with the bounding box area are scanned in the multi-level feature map, and the attention weights are dynamically adjusted according to the contrast relationship between the pixels inside and outside the bounding box. The decoder outputs a multi-channel probability map P by gradually upsampling the global semantics and local details. t (u, v), and finally convert the candidate mask with the highest probability (usually the result with the highest spatial overlap with the input box and the best edge continuity) into a binary mask (i.e., segmentation mask), where is the indicative function, and τ is the decision threshold. This process relies on densely supervised learning of the box-mask relationship during pre-training, enabling the model to quickly associate 2D geometric constraints with pixel-level segmentation results during inference without the need for additional fine-tuning. During deployment, NVIDIA's open-source NanoSAM project was used. The lightweight ViT image encoder of the MobileSAM model was used as the teacher model, and the ResNet18 model was used as the student model. The image encoder of the NanoSAM model was distilled. The trained image encoder and mask decoder were then converted to FP16 format to accelerate detection efficiency. The model was then exported using the ONNX universal format for neural network models. TensorRT was used to build the model engine to accelerate on-device model inference.

[0067] Step 2: Perform point cloud post-processing based on the ordered point cloud data of the current frame and the obstacle image segmentation mask.

[0068] Get the colored ordered point cloud P of the current frame t and the previous image segmentation mask M t , use the point-by-point traversal method to segment the point cloud data (i.e. point cloud clusters) corresponding to the obstacles P o ={p i ∈P t |M t (x i ,y i )∈l o}, where p i The mask corresponds to a single point cloud data point, M t (x i ,y i ) is the mask data point at the corresponding coordinate.

[0069] Specifically obtain point cloud data P oProcess: For static and dynamic obstacle categories, the descriptor method based on the moment of inertia and eccentricity of the PCL (PointCloud Library) point cloud library is used to extract the oriented bounding box (OBB) corresponding to the obstacle so that the subsequent planning system can perform obstacle avoidance processing. When using PCL to process point cloud data, the calculation process of the oriented bounding box mainly includes the following steps: First, the point cloud data is pre-processed to filter out ground points and non-obstacle noise points, and the local point cloud clusters of the target (i.e., point cloud clusters) are extracted. Then the first-order statistics of the point cloud clusters, i.e., the centroid, are calculated. As the spatial reference point of the bounding box, where i is the point cloud data point index and N is the total number of point cloud data points; the core second-order statistic covariance matrix It will be used to analyze the distribution characteristics of the point cloud, which characterizes the distribution of the moment of inertia of the point cloud in different axes. By performing eigendecomposition on ∑, three orthogonal eigenvectors v1, v2, v3 and their corresponding eigenvalues ​​λ1≥λ2≥λ3 are obtained. The main eigenvector v1 corresponding to the largest eigenvalue defines the main extension direction of the point cloud, which is used as the direction reference axis of the bounding box. In order to determine whether rotation alignment is required, the target eccentricity index needs to be further calculated. The eccentricity is defined as When it exceeds the set threshold, it indicates that the point cloud has significant directional differences. At this time, the main direction v1 is used as the long axis of the bounding box, and the point cloud is projected onto the planes of each main axis in this rotating coordinate system. The size parameters (length, width, and height) of the bounding box are determined by finding the extreme points of the projection. The resulting six-degree-of-freedom bounding box (i.e., an oriented bounding box) with its center at μ and its rotation axis defined by the eigenvectors can provide accurate obstacle space occupancy constraints for the motion planning system. The entire process improves the robustness of obstacle representation in complex environments by fusing geometric statistics with motion features.

[0070] For semi-dynamic obstacles, an efficient plane fitting and segmentation method based on the cuPCL (CUDA Point Cloud Library) point cloud library is used to determine the plane normal direction to obtain the obstacle orientation. This allows for subsequent detection of the angle between targets such as doors, windows, nets, curtains, and walls, and determines the operating strategy of the unmanned aerial platform.

[0071] Detailed description of semi-dynamic obstacle orientation vector extraction: First, the original point cloud is downsampled and outliers are removed. A filter based on voxel grid downsampling (VoxelGrid) is used and the cuPCL parallel computing architecture is used to accelerate data preprocessing, reducing the amount of data for subsequent point cloud reading and processing while ensuring the accuracy of the point cloud data. Then, the random sampling consensus (RANSAC) algorithm based on the cuPCL library is used to fit the plane and segmentation method, and the model fitting coefficient is used as the plane normal coefficient πo :a o x+b o y+c o z+d o =0, and then obtain the obstacle direction vector The core of the RANSAC algorithm is to iteratively search for the best plane model: each time randomly select three non-collinear points to construct a candidate plane equation a o x+b o y+c o z+d o = 0, calculate the projection distance of all points to the plane, and count the number of inliers that meet the threshold condition. After multiple rounds of iteration, the plane with the highest inlier coverage is selected as the final model, and the plane parameter a is optimized by the least squares method based on the coordinates of all inliers. o ,b o ,c o , so that |a o x i +b o y i +c o z i +d o The residual sum of squares is minimized, and the normalization constraint is passed Determine the normal vector direction.

[0072] Step 3: Based on the segmented point cloud and the fitted plane information obtained through post-processing, the door and window opening and closing angles are detected and the semi-dynamic obstacle status information is extracted.

[0073] Obtain the colored ordered point cloud of the current frame, remove the point cloud data corresponding to the obstacle, segment the point cloud data of the background part, and use the fitting plane and segmentation method based on the cuPCL library to determine the normal direction of the wall plane to obtain the wall orientation vector. Finally, combine it with the previously obtained obstacle orientation vector and determine the relative relationship and angle between the door and window and the wall according to the vector transformation rules.

[0074] Detailed description of the wall orientation vector extraction: First, the cuPCL library-based random sampling consensus (RANSAC) algorithm is used to fit the plane and segmentation method, and the model fitting coefficient is used as the plane normal coefficient π w :a w x+b w y+c w z+d w =0, and then obtain the wall direction vector The RANSAC algorithm process is as described above.

[0075] Then based on the obstacle direction vector and the wall direction vector Get the quaternion q and Euler angles (φ, θ, ψ) corresponding to the vector transformation, and use the relative positions of the camera and the obstacle on the wall and the range of Euler angles [-π / 2, π / 2] to regularly determine the inside and outside opening and closing states S of obstacles such as doors and windows. o and opening and closing angle θ o , or the connection status of obstacles such as nets and curtains in, is the normal vector, is the mean vector of the normal vectors, and ||·|| represents the distance metric.

[0076] Detailed description of multi-dimensional feature extraction of semi-dynamic obstacles: For rigid semi-dynamic obstacles such as doors and windows, after obtaining the obstacle direction vector and the wall orientation vector After that, the vector direction is calibrated first: if the dot product of the two vectors is negative (i.e. the angle is greater than 90 degrees), the obstacle direction vector is reversed to ensure that the two vectors are in the same half space. Then the rotation transformation relationship between the two is solved through geometric transformation: based on the construction principle of quaternion, the quaternion q can be generated by determining the minimum angle rotation axis and the rotation angle. This quaternion describes the relative rotation from the wall direction vector to the obstacle direction vector; specifically, the rotation axis is the cross product of the two vectors. After normalization, the rotation angle θ is determined by the arc cosine formula θ = The final quaternion is defined as q=cos(θ / 2)+(sin(θ / 2)(u xi +u yi +u zk )), where (u x ,u y ,u z ) is the unit vector of the rotation axis. The quaternion q is then converted into a three-dimensional rotation matrix R by expanding the quaternion multiplication rule. The matrix contains the rotation components around the three axes of the coordinate system. Based on the rotation matrix, the Euler angle decomposition rule of ZYX order is used to extract the final Euler angle triplet and the opening and closing angle θ o To prevent the angle from exceeding the range of [-π / 2,π / 2], the calculation results need to be normalized: if the absolute value of the angle exceeds π / 2, it is mapped to the valid range through phase adjustment. o The core of the calculation is the discrete characteristics of the normal vector after the sub-region segmentation. First, the curtain target is divided into N sub-regions through spatial block processing, and the normal vector is independently performed on each sub-region using the RANSAC method described above. (where k = 1, 2, ..., N) extraction operations, combined with cuPCL's parallel segmentation capabilities, ensure real-time processing efficiency of high-frequency data. Next, calculate the average vector of all normal vectors Then, the square of the Euclidean distance between the normal vector of each sub-region and the average vector is calculated point by point. And take the mean to obtain the overall dispersion index If the connection state of each sub-region of the curtain is stable (such as taut, fixed), the local normal vector distribution will be similar, C o tends to a smaller value; if there is local relaxation or fracture (such as rupture, flutter), the local normal vectors of different regions are significantly different, resulting in C o This indicator is embedded in the decision logic after weighted screening, providing a quantitative basis for the flight platform to judge obstacle stability, plan obstacle avoidance paths, or determine the threshold for intervention in operations.

[0077] Step 4: Based on the extracted multi-dimensional features of obstacles, perform passability judgment and path planning.

[0078] The extracted multi-dimensional features of dynamic, static and semi-dynamic obstacles (i.e., state information of semi-dynamic obstacles and directional bounding boxes of non-semi-dynamic obstacles) are obtained to make passability judgment.

[0079] For static and dynamic obstacle categories, traditional obstacle avoidance algorithms such as RRT and A* are used to avoid them based on the identified distance and directional bounding box.

[0080] For semi-dynamic obstacles, feasible interaction points are determined based on the identified multi-dimensional features and other factors such as the platform's dimensions, allowing for navigation. The latter utilizes a trajectory generation algorithm that incorporates interaction points and preview paths, dynamically adjusting trajectory generation parameters based on obstacle characteristics to achieve dynamic adaptation to the environment. The generated trajectory is then input into the flight controller, enabling autonomous navigation by the unmanned aerial vehicle.

[0081] The trajectory generation algorithm mainly includes the following steps: First, based on the inner and outer opening and closing states and opening and closing angles of the rigid semi-dynamic obstacle obtained in step 3 and the connection state of the flexible semi-dynamic obstacle, a passability decision is made: when the rigid semi-dynamic obstacle is in the inner or outer opening state and the opening and closing angle θ o When the flight platform is allowed to pass through (the opening and closing angle is too large), or the connection state index of the flexible semi-dynamic obstacle is lower than the stability threshold, indicating that its structure is tight, the algorithm will generate a passing obstacle avoidance trajectory; when the rigid semi-dynamic obstacle is in the inward opening state and the opening and closing angle θ oWhen the flight platform is allowed to interact (the opening and closing angle is small and non-zero), or the continuous state index of the flexible semi-dynamic obstacle is higher than the stability threshold, indicating that its structure is loose, the algorithm will generate an interactive operation trajectory. For the above scenarios, the trajectory generation adopts a minimal snap optimization framework based on time segmentation, the goal is to minimize the fourth-order derivative (snap) energy integral of the trajectory, that is, to construct the objective function Among them, p i (t) is the i-th polynomial trajectory, t∈[t i-1 ,t i ] is the time interval, M is the total number of segments, d 4 is the symbol for the fourth-order derivative. This optimization process must satisfy multiple constraints, including position, velocity, and acceleration continuity constraints at the starting point and the preview path points required in interactive scenarios, as well as flight platform dynamics limitations (such as maximum velocity, normal acceleration, and attitude angular rate thresholds). Specifically, in interactive operation scenarios, the trajectory is constructed as a three-segment structure: the initial segment runs from the current position to the point of contact with the obstacle surface, the middle segment penetrates the obstacle boundary to the inside, and the final segment extends to a preview point at a fixed distance behind the obstacle, such as 1 meter, to ensure that stable adjustment space is reserved for the flight platform after crossing. The connection points of each segment must meet the continuity of position, velocity, and acceleration, and the end constraints of zero velocity and zero acceleration are applied at the preview point to achieve smooth dwell. The resulting minimal snap trajectory is input into the underlying flight controller, which uses the control algorithm to achieve precise tracking in a dynamic environment, thereby completing the navigation task.

[0082] Finally, the overall architecture is optimized and deployed. The algorithm of the above process is written into a ROS (Robotics Operating System) software package containing four nodes: image processing, point cloud processing, obstacle feature extraction, and trajectory planning. In response to the multi-machine deployment requirements across distributions, a container construction pipeline is designed to facilitate the rapid deployment of subsequent real-machine algorithms. For heterogeneous embedded platforms on the end side, the container construction pipeline uses Docker container technology to unify the packaging of runtime environments such as ROS, Python, and TensorRT, achieve environmental standardization and rapid deployment, reduce configuration and compatibility issues, improve the overall efficiency and stability of the system, and support flexible migration and expansion of different hardware platforms.

[0083] The perception system of the semi-dynamic obstacle recognition method for the unmanned aerial platform is deployed on the flight platform and is loaded with an Nvidia OrinNX 16GB high-performance embedded platform, including an Arm Cortex-A78AE microprocessor (CPU) and an Ampere graphics processor (GPU), as well as an Intel RealSense D435i depth camera and a Leixun Nora+ flight controller. Figure 2The unmanned aerial platform equipment structure diagram is shown in the figure. Figure 3 This is a diagram of the software package architecture provided in the embodiment of this application. Figure 4 Flowchart of the real-time detection and planning method for multi-dimensional features of obstacles provided in an embodiment of the present application.

[0084] Example 3

[0085] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store processing data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a semi-dynamic obstacle recognition method for an unmanned aerial platform is implemented.

[0086] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0087] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0088] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0089] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0090] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0091] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0092] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0093] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0094] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A semi-dynamic obstacle recognition method for an unmanned aerial platform, characterized in that: include: Obtain high-resolution color images and corresponding ordered point cloud data; performing open set object detection on the color image to obtain a bounding box of the obstacle, and performing visual segmentation on the color image and the bounding box to obtain a segmentation mask of the obstacle; the obstacle includes a non-semi-dynamic obstacle and a semi-dynamic obstacle; the non-semi-dynamic obstacle is a static obstacle or a dynamic obstacle, or the non-semi-dynamic obstacle includes a static obstacle and a dynamic obstacle; Obtaining point cloud clusters corresponding to obstacles based on the ordered point cloud data and the segmentation mask; When the obstacles include both non-semi-dynamic and semi-dynamic obstacles, an oriented bounding box is calculated for the point cloud clusters of the non-semi-dynamic obstacles, and a heading vector is fitted for the point cloud clusters of the semi-dynamic obstacles. Then, based on the heading vectors of the semi-dynamic obstacles and the obtained heading vectors of the walls, the semi-dynamic obstacle state information is obtained. Finally, based on the semi-dynamic obstacle state information and the oriented bounding boxes of the non-semi-dynamic obstacles, the passability of all obstacles is determined to obtain the planned trajectory of the unmanned aerial platform.

2. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: The open set object detection is performed using an optimization model based on OWL-ViT. Open set object detection is performed on the color image to obtain a bounding box of the obstacle, specifically including: Dividing the color image into image blocks, embedding the image blocks into vector sequences using a linear mapping, and encoding the vector sequences using a multi-layer visual transformer to obtain a global feature map; Use the text encoder to transform the obtained obstacle category to obtain the text feature vector; Performing cross-modal feature alignment on the global feature map and the text feature vector using a cross-attention mechanism to obtain aligned multimodal features; A candidate bounding box is generated based on the multimodal features, and the candidate bounding box is filtered using non-maximum suppression to output a bounding box of the obstacle.

3. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: The visual segmentation is performed using a model based on SAM project distillation; the model based on SAM project distillation includes an encoder and a decoder; the encoder includes a visual backbone network and a hint encoding module; Performing visual segmentation on the color image and the bounding box to obtain a segmentation mask of the obstacle specifically includes: Inputting the color image into the visual backbone network to obtain a multi-scale feature map; Associating the bounding box with the multi-scale feature map using a hint encoding module to obtain an association result; The decoder is used to upsample the association result to obtain a multi-channel probability map, and the multi-channel probability map is converted into a segmentation mask of the obstacle.

4. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: Based on the ordered point cloud data and the segmentation mask, a point cloud cluster corresponding to the obstacle is obtained, specifically including: According to the segmentation mask, the ordered point cloud data is processed using a point-by-point traversal method to obtain point cloud clusters corresponding to obstacles.

5. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: Compute an oriented bounding box for a point cloud cluster of non-semi-dynamic obstacles, specifically: Performing point cloud preprocessing on the point cloud cluster of the non-semi-dynamic obstacle to obtain a first point cloud preprocessing result; the point cloud preprocessing includes filtering denoising and clustering segmentation; Performing first-order statistics calculation on the first point cloud preprocessing result to obtain centroid coordinates, and performing second-order statistics calculation on the centroid coordinates to obtain a covariance matrix; Perform eigendecomposition on the covariance matrix to obtain a main direction vector, and project the first point cloud preprocessing result onto the main direction vector to obtain an oriented bounding box.

6. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: Fitting the orientation vector of the point cloud cluster of the semi-dynamic obstacle includes: Performing accelerated point cloud preprocessing on the point cloud cluster of the semi-dynamic obstacle to obtain a second point cloud preprocessing result; the accelerated point cloud preprocessing includes downsampling and outlier removal; The RANSAC algorithm is used to perform plane fitting and normal extraction on the second point cloud preprocessing result; and the orientation vector of the semi-dynamic obstacle is obtained.

7. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: The semi-dynamic obstacle includes a rigid semi-dynamic obstacle and / or a flexible semi-dynamic obstacle; the semi-dynamic obstacle state information includes the opening and closing angle of the rigid semi-dynamic obstacle and / or the connection state of the flexible semi-dynamic obstacle; based on the orientation vector of the semi-dynamic obstacle and the obtained orientation vector of the wall, the semi-dynamic obstacle state information is obtained, specifically including: removing obstacle point cloud data from the ordered point cloud data to obtain wall point cloud data; Using the RANSAC algorithm to perform plane fitting and normal extraction on the wall point cloud data to obtain the orientation vector of the wall; When the semi-dynamic obstacle is a rigid semi-dynamic obstacle, the orientation vector of the rigid semi-dynamic obstacle and the orientation vector of the corresponding wall are calculated using the quaternion method to obtain a quaternion, and the quaternion is subjected to Euler angle decomposition and opening and closing angle determination to obtain the opening and closing angle of the rigid semi-dynamic obstacle; When the semi-dynamic obstacle is a flexible semi-dynamic obstacle, performing an average normal vector calculation on the orientation vector of the flexible semi-dynamic obstacle, and performing a dispersion index calculation on the flexible semi-dynamic obstacle based on the orientation vector of the flexible semi-dynamic obstacle and the average normal vector to obtain a connection state of the flexible semi-dynamic obstacle; When the semi-dynamic obstacle includes a rigid semi-dynamic obstacle and a flexible semi-dynamic obstacle, the quaternion method is used to calculate the orientation vector of the rigid semi-dynamic obstacle and the orientation vector of the corresponding wall to obtain a quaternion. The quaternion is then subjected to Euler angle decomposition and opening and closing angle determination to obtain the opening and closing angles of the rigid semi-dynamic obstacle. In addition, the average normal vector of the orientation vector of the flexible semi-dynamic obstacle is calculated. Based on the orientation vector of the flexible semi-dynamic obstacle and the average normal vector, a discrete index of the flexible semi-dynamic obstacle is calculated to obtain the connection status of the flexible semi-dynamic obstacle.

8. The semi-dynamic obstacle recognition method for an unmanned aerial platform according to claim 1, characterized in that: Based on the semi-dynamic obstacle state information and the oriented bounding boxes of the non-semi-dynamic obstacles, a passability judgment is performed on all obstacles to obtain a planned trajectory of the unmanned aerial platform, specifically including: Based on the oriented bounding box of the non-semi-dynamic obstacle, the first trajectory is generated using the RRT algorithm or the A* algorithm; Based on the semi-dynamic obstacle state information, a second trajectory is generated using a trajectory generation algorithm of interaction points and preview paths; Optimizing the second trajectory using Minimal Snap trajectory optimization to obtain a third trajectory; A planned trajectory is determined based on the first trajectory and the third trajectory.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the semi-dynamic obstacle recognition method for an unmanned aerial platform according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the semi-dynamic obstacle recognition method for an unmanned aerial platform according to any one of claims 1 to 8 is implemented.