Scene recognition and incremental mapping method based on scene topology graph perceived by unmanned aerial vehicle
Patent Information
- Application Number
- CN202611142803.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-30
AI Technical Summary
这类方法虽然能够实现一定程度的定位和建图,但在大范围、长时间飞行任务中通常需要保存大量关键帧图像、局部点云或高密度三维地图数据,导致历史数据库规模不断增大,增加机载存储压力和匹配计算量,难以满足无人机平台对轻量化和实时性的要求
[0040]1、本方法通过筛选出关键帧,并从关键帧中抽取全局外观描述子和稳定目标,通过场景节点、地标节点以及各节点之间的关系边构建各关键帧的拓扑图,在减少完整图像帧和三维点云数据存储需求的同时,保留场景外观、目标语义和空间信息,并且结合拓扑图进行匹配,能够降低单一特征匹配导致的误识别风险,提高无人机场景识别的准确性和鲁棒性;
Smart Images

Figure CN122695522B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous perception and spatial intelligence technology for unmanned aerial vehicles (UAVs), specifically relating to a method for scene recognition and incremental mapping based on the topology map of the scene perceived by the UAV. Background Technology
[0002] With the increasing demand for applications such as autonomous inspection, emergency search and rescue, environmental monitoring, and intelligent navigation, drones need to continuously perceive the scene in complex environments and determine whether the current scene is a previously known scene in order to support autonomous localization, path planning, and incremental mapping. Especially in semantic navigation or visual-language navigation tasks, drones not only need to acquire the spatial structure of the environment, but also need to identify target instances in the scene and their relative positional relationships, so as to associate the perception results with the target objects and orientation relationships in the task instructions.
[0003] Existing autonomous navigation and mapping methods for unmanned aerial vehicles (UAVs) largely rely on simultaneous localization and mapping (SLAM) technology, describing the environment through image feature points, line features, or dense point cloud data. While these methods can achieve a certain level of localization and mapping, large-scale, long-duration flight missions typically require storing a large amount of keyframe images, local point clouds, or high-density 3D map data. This leads to a continuous increase in the size of the historical database, increasing onboard storage pressure and matching computational load, making it difficult to meet the lightweight and real-time requirements of UAV platforms. Furthermore, traditional map representation mainly focuses on low-level geometric structures, lacking effective representation of stable semantic targets and their spatial relationships. In inspection or semantic navigation tasks, systems often need to understand high-level semantic information such as target objects, their locations, and relative relationships between targets. Relying solely on low-level geometric features or dense point clouds cannot directly support the expression of such semantic relationships and is also not conducive to subsequent association with natural language commands or high-level mission planning.
[0004] Furthermore, existing scene recognition methods are susceptible to changes in viewpoint, illumination, partial occlusion, and repetitive scene structures. Relying solely on image appearance features may lead to mismatches between similar scenes, while relying solely on geometric features is easily affected by sparse point clouds or noise, resulting in insufficient accuracy and robustness in scene recognition. Summary of the Invention
[0005] The purpose of this invention is to address the problems mentioned in the background art by proposing a scene recognition and incremental mapping method based on the scene topology map perceived by UAVs.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] This invention proposes a scene recognition and incremental mapping method based on UAV-perceived scene topology maps, comprising:
[0008] Acquire image frames, 3D point cloud data, and UAV pose data continuously collected by the UAV within the current preset time period for the same target scene;
[0009] Filter consecutive image frames and select those that meet preset conditions as keyframes;
[0010] Scene extraction and target detection are performed on each keyframe to obtain the global appearance descriptor, target region, semantic category and confidence of each keyframe, and the three-dimensional centroid coordinates of the target are calculated using the three-dimensional point cloud corresponding to the target region.
[0011] The stability score of a target is calculated using the target's confidence level, and targets with a stability score that reaches the first threshold are considered stable targets, with each stable target serving as a landmark node.
[0012] Construct a topology graph for each keyframe, which includes scene nodes, landmark nodes, and the relationship edges between nodes. The scene nodes contain the global appearance descriptor of the current keyframe and the corresponding UAV pose data, while the landmark nodes contain the semantic category, confidence, 3D centroid coordinates, and stability score of the corresponding stable target.
[0013] Each topology graph is matched with the topology subgraphs in the historical database. The comprehensive scene similarity between each topology graph and the topology subgraphs in the historical database is calculated. For each topology graph, the maximum comprehensive scene similarity is compared with a preset similarity value. If it is greater than the preset similarity value, the target scene corresponding to the current topology graph is a known scene; otherwise, it is an unknown scene. The current topology graph is then stored as a new topology subgraph in the historical database, thereby realizing incremental graph building on the historical database.
[0014] Preferably, the UAV pose data includes the UAV's position and attitude, and the attitude includes the heading angle and pitch angle.
[0015] Preferably, the step of filtering consecutive image frames and selecting image frames that meet preset conditions as keyframes includes:
[0016] Step 1: Use the first image frame in a series of consecutive image frames as the initial keyframe;
[0017] Step II: Each image frame following the current keyframe is called a candidate image frame, and the first score between each candidate image frame and the current keyframe is calculated in turn. If the first score is greater than the second threshold, the current candidate image frame is used as the new current keyframe; otherwise, the current candidate image frame is not used as a keyframe.
[0018] Step III: Repeat Step II and iterate continuously until all image frames in a series of consecutive image frames have participated in the calculation of the first score at least once. Then stop the iteration and complete the screening of key frames in the series of consecutive image frames.
[0019] The first score is the weighted sum of the changes in UAV position, UAV attitude, and time between the candidate image frame and the current keyframe.
[0020] Preferably, when performing scene extraction and target detection on each key frame, the key frame is processed using an image processing model to obtain the global appearance descriptor of the current key frame, and the key frame is processed using a target detection model to obtain the target region, semantic category and confidence of each target in the current key frame.
[0021] When using the three-dimensional point cloud data corresponding to the target area to calculate the three-dimensional centroid coordinates of the target, the three-dimensional point cloud data is projected into the image frame. The three-dimensional point cloud falling into the target area is taken as the point cloud of the target, and the average coordinate of all the point clouds of the target is taken as the three-dimensional centroid coordinates of the target.
[0022] Preferably, when calculating the stability score of the target, the confidence level of the target is weighted and summed with the number of observations of the target in adjacent frames to obtain the stability score;
[0023] The number of times the target is observed in adjacent frames is the sum of the number of times the target appears in the current keyframe, the previous keyframe, and the next keyframe.
[0024] Preferably, the relationship edges of each node in the topology graph of each keyframe include the observation relationship edges between scene nodes and each landmark node, as well as the spatial relationship edges between each landmark node;
[0025] The observation relationship edge includes the relative first spatial distance, relative first azimuth angle, and relative first pitch angle between the UAV in the scene node and the stable target corresponding to the landmark node. The relative first spatial distance, relative first azimuth angle, and relative first pitch angle are all calculated by the position of the UAV in the scene node and the three-dimensional centroid coordinates of the stable target corresponding to the landmark node.
[0026] The spatial relationship edge includes the relative second spatial distance between two landmark nodes, and the relative second spatial distance is calculated using the three-dimensional centroid coordinates of the stable targets corresponding to the two landmark nodes.
[0027] Preferably, the step of matching each topology graph with the topology subgraphs in the historical database and calculating the comprehensive scene similarity between each topology graph and the topology subgraphs in the historical database includes:
[0028] The similarity between the global appearance descriptor in the current topology graph and the global appearance descriptors of all topology subgraphs in the historical database is calculated, and the top K most similar topology subgraphs are selected from the historical database to form a candidate matching set.
[0029] Each landmark node in the current topology graph is called the first node, and each landmark node in each topology subgraph in the candidate matching set is called the second node.
[0030] For each topological subgraph in the candidate matching set, each first node is matched with all second nodes in the current topological subgraph. When the semantic categories of the first node and the second node are the same and their stability scores differ within a first preset range, the first node and the second node are called a matching landmark node pair.
[0031] Based on all matching landmark node pairs between the current topological graph and the current topological subgraph, calculate the semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph;
[0032] The semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph are weighted and summed to obtain the comprehensive scene similarity between the current topological graph and the current topological subgraph.
[0033] Preferably, the step of calculating the semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph based on all matching landmark node pairs between the current topological graph and the current topological subgraph includes:
[0034] Calculate the matching score of the first node in each matching landmark node pair, and the matching score is equal to the weighted sum of the confidence and stability scores of the first node in the current matching landmark node pair;
[0035] The semantic similarity score between the current topological graph and the current topological subgraph is obtained by weighted averaging of the matching scores of all first nodes in the current topological graph.
[0036] For each matching landmark node pair, compare the observation relationship edge of the first node and the observation relationship edge of the second node. If the relative second spatial distance, relative second azimuth angle and relative second elevation angle of the first node and the second node are respectively within the corresponding second preset range, then the current matching landmark node pair is called the geometric interior point; otherwise, it is called the geometric exterior point. The relative second azimuth angle and relative second elevation angle are calculated by the three-dimensional centroid coordinates of the stable targets corresponding to the first node and the second node, respectively.
[0037] Calculate the proportion of geometric interior points in all matching landmark node pairs, and use this proportion as the geometric similarity score between the current topology graph and the current topology subgraph.
[0038] Preferably, each topological subgraph in the historical database includes scene nodes, landmark nodes, and the relationship edges between the nodes.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] 1. This method selects keyframes and extracts global appearance descriptors and stable targets from them. It constructs a topology graph of each keyframe by using scene nodes, landmark nodes, and the relationship edges between nodes. This reduces the need for storing complete image frames and 3D point cloud data while preserving scene appearance, target semantics, and spatial information. Furthermore, by combining the topology graph for matching, it can reduce the risk of misidentification caused by single feature matching and improve the accuracy and robustness of UAV scene recognition.
[0041] 2. For the topology graph of unknown scenarios that are not matched, it is stored as a new topology subgraph in the historical database to realize incremental graph building of the historical database. Attached Figure Description
[0042] Figure 1 This is a flowchart illustrating the scene recognition and incremental mapping method based on the scene topology map perceived by the UAV according to the present invention.
[0043] Figure 2 This is a schematic diagram of the topology graph in this invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention.
[0046] In one embodiment, such as Figures 1-2 As shown, a method for scene recognition and incremental mapping based on UAV-perceived scene topology maps is provided, including:
[0047] Step 1: Acquire image frames, 3D point cloud data, and UAV pose data continuously collected by the UAV for the same target scene within the current preset time period (e.g., within 3 seconds up to the current moment);
[0048] It should be noted that during the autonomous flight of the UAV, the target scene is continuously observed by the RGB camera, 3D LiDAR, and inertial measurement unit on board the UAV, acquiring continuous image frames, 3D point cloud data, and UAV pose data of the target scene. Among them, the RGB camera is used to collect image frames of the target scene; the 3D LiDAR is used to collect 3D point cloud data of the target scene, which is used to describe the spatial distribution of objects in the target scene; the inertial measurement unit is used to collect acceleration and angular velocity information of the UAV during flight, and combines it with the 3D LiDAR point cloud matching results for fusion positioning to obtain the UAV's pose data.
[0049] The UAV pose data includes the UAV's position and attitude. Position is used to characterize the UAV's location in the spatial coordinate system, and attitude is used to characterize the UAV's observation direction relative to the target scene. The attitude includes the heading angle and pitch angle.
[0050] During the acquisition process, the RGB camera, 3D LiDAR, and inertial measurement unit record data according to a unified time reference, so that the image frames, 3D point cloud data, and UAV pose data at the same observation time can correspond to each other. For each observation time, the corresponding image frames, 3D point cloud data, and UAV pose data are used as a set of target scene observation data. This set of target scene observation data is used for subsequent keyframe screening, scene extraction, target detection, and topology map construction.
[0051] Before the drone flies, the installation position relationship between the RGB camera and the 3D LiDAR is jointly calibrated to obtain the spatial transformation relationship between the RGB camera coordinate system and the 3D LiDAR coordinate system. Based on this spatial transformation relationship, the 3D point cloud collected by the 3D LiDAR can be transformed into the RGB camera coordinate system, and the transformed 3D point cloud can be projected onto the RGB image plane, thereby establishing the correspondence between the target area in the RGB image and the 3D point cloud points.
[0052] Step 2: Filter consecutive image frames, selecting those that meet preset conditions as keyframes, including:
[0053] Step 2.1: Use the first image frame in a series of consecutive image frames as the initial keyframe;
[0054] Step 2.2: Each image frame following the current keyframe is called a candidate image frame, and the first score between each candidate image frame and the current keyframe is calculated sequentially. When the first score is greater than the second threshold (i.e., the preset condition is met, and the second threshold is 0.4), the current candidate image frame is used as the new current keyframe; otherwise, the current candidate image frame is not used as a keyframe.
[0055] Step 2.3: Repeat step 2.2 and iterate continuously until all image frames in a series of consecutive image frames have participated in the calculation of the first score at least once. Then stop the iteration and complete the screening of key frames in the series of consecutive image frames.
[0056] The first score is the weighted sum of the changes in UAV position, UAV attitude, and time between the candidate image frame and the current keyframe. The corresponding calculation formula is as follows:
[0057] ;
[0058] in, For each candidate image frame and the current The first score between keyframes at any given moment. For each candidate image frame and the current Changes in drone position between keyframes at different times For each candidate image frame and the current The amount of drone attitude change between keyframes at any given moment. For each candidate image frame and the current The amount of time change between keyframes at any given moment. , and They are respectively , and The weighting coefficients (adjusted according to the actual application scenario).
[0059] Step 3: Perform scene extraction and target detection on each keyframe to obtain the global appearance descriptor, target region, semantic category, and confidence score of each keyframe. Then, use the 3D points corresponding to the target region to calculate the 3D centroid coordinates of the target, including:
[0060] Step 3.1: When performing scene extraction and object detection on each keyframe, use an image processing model (such as MobileNet or ResNet, with the keyframe as the model input) to process the keyframe and obtain the global appearance descriptor of the current keyframe. Use an object detection model (such as YOLO series or Faster R-CNN, with the keyframe as the model input) to process the keyframe and obtain the target region, semantic category, and confidence of each object in the current keyframe. Both the image processing model and the object detection model are pre-trained and do not require additional training.
[0061] Step 3.2: When calculating the 3D centroid coordinates of the target using the 3D point cloud data corresponding to the target region, the 3D point cloud data is projected onto the image frame. The 3D point cloud falling into the target region is taken as the target's point cloud, and the average coordinate of all the target's point clouds is taken as the target's 3D centroid coordinates. The corresponding calculation formula is as follows:
[0062] ;
[0063] in, For the first keyframe The three-dimensional centroid coordinates of the target. For the first keyframe The first target corresponds to the The coordinates of a point cloud, For the first keyframe The number of all point clouds corresponding to each target.
[0064] Step 4: Calculate the stability score of the target using the confidence score of the target, and designate the targets whose stability scores reach the first threshold (e.g., the first threshold is 0.6) as stable targets, and each stable target as a landmark node;
[0065] In calculating the target's stability score, the target's confidence level is weighted and summed with the number of observations of the target in adjacent frames to obtain the stability score. The corresponding calculation formula is as follows:
[0066] ;
[0067] in, For the first keyframe Stability score of each target, For the first keyframe Confidence level of each target For the first keyframe The number of times a target is observed in adjacent frames. and They are respectively and The weighting coefficients (set according to the actual scenario) are as follows: the number of times the target is observed in adjacent frames is the sum of the number of times the target appears in the current keyframe, the previous keyframe, and the next keyframe; the first threshold is expressed as... ,when Then the current number The goal is a stable goal.
[0068] Step 5: Construct a topology graph for each keyframe, including a scene node (one), multiple landmark nodes, and edges connecting the nodes. The scene node contains the global appearance descriptor of the current keyframe and the corresponding UAV pose data, while the landmark nodes contain the semantic category, confidence score, 3D centroid coordinates, and stability score of the corresponding stable target. Figure 2 The image shown is an example of a topology graph, containing one scene node and four landmark nodes (in...). Figure 2 The middle represents a landmark node. Landmark nodes Landmark nodes and landmark nodes ), and the relationship edges of each node.
[0069] Among them, the relationship edges of each node in the topology graph of each keyframe include the observation relationship edges between scene nodes and each landmark node, as well as the spatial relationship edges between each landmark node;
[0070] The observation relationship edge includes the relative first spatial distance, relative first azimuth angle, and relative first pitch angle between the UAV and the stable target corresponding to the landmark node in the scene node. The relative first spatial distance, relative first azimuth angle, and relative first pitch angle are all calculated by the position of the UAV in the scene node and the three-dimensional centroid coordinates of the stable target corresponding to the landmark node (that is, the first spatial distance is calculated using the position coordinates of the UAV and the three-dimensional centroid coordinates, and the relative first azimuth angle and relative first pitch angle are calculated using the position coordinates of the UAV and the three-dimensional centroid coordinates).
[0071] The spatial relationship edge includes the relative second spatial distance between two landmark nodes, and the relative second spatial distance is calculated by the three-dimensional centroid coordinates of the stable targets corresponding to the two landmark nodes (i.e., calculating the spatial distance between the two three-dimensional centroid coordinates).
[0072] The topology of each keyframe is represented as follows:
[0073] ;
[0074] in,
[0075] ;
[0076] in, For the present Topology of keyframes at different times. For the present The set of nodes in the topology graph of a time-key frame. For the present The set of relational edges in the topological graph of keyframes at any given time. For the present Scene nodes in the topology graph of time-key frames. For the present The topology of the keyframe at time step A landmark node, For the present Scene nodes in the topology graph of time-key frames and the first The observation relationship edges between each landmark node For the present The topology of the keyframe at time step The first landmark node and the first The spatial relationship edges of the landmark nodes, among which , Belongs to the present The set of all landmark nodes in the topology graph of a time-key frame.
[0077] Step 6: Match each topology graph with the topology subgraphs in the historical database, calculate the comprehensive scene similarity between each topology graph and the topology subgraphs in the historical database. For each topology graph, compare the maximum comprehensive scene similarity with a preset similarity value (e.g., 0.75). If the value is greater than the preset similarity value, the target scene corresponding to the current topology graph is a known scene; otherwise, it is an unknown scene. The current topology graph is then stored as a new topology subgraph in the historical database, achieving incremental graph building on the historical database, including:
[0078] Step 6.1: Match each topology graph with the topology subgraphs in the historical database, and calculate the comprehensive scene similarity between each topology graph and the topology subgraphs in the historical database, including:
[0079] Step 6.1.1: Calculate the (cosine) similarity between the global appearance descriptor of the current topology graph and the global appearance descriptors of all topology subgraphs in the historical database, and select the top K topology subgraphs with the highest similarity from the historical database (e.g., K is 5) to form a candidate matching set;
[0080] Step 6.1.2: Refer to each landmark node in the current topology graph as the first node, and refer to each landmark node in each topology subgraph in the candidate matching set as the second node;
[0081] Step 6.1.3: For each topological subgraph in the candidate matching set, match each first node with all second nodes in the current topological subgraph. When the semantic categories of the first node and the second node are the same and their stability scores differ within a first preset range (e.g., the first preset range is within 0.2), then the first node and the second node are called a matching landmark node pair.
[0082] Step 6.1.4: Based on all matching landmark node pairs between the current topological graph and the current topological subgraph, calculate the semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph.
[0083] Step 6.1.4.1: Calculate the matching score of the first node in each matching landmark node pair (calculate the matching landmark node pairs between the current topology graph and the current topology subgraph), and the matching score is equal to the weighted sum of the confidence and stability scores of the first node in the current matching landmark node pair (for the first and second nodes with inconsistent semantic categories, the matching score is equal to zero).
[0084] Step 6.1.4.2: Calculate the weighted average of the matching scores of all first nodes in the current topology graph to obtain the semantic similarity score between the current topology graph and the current topology subgraph. The corresponding calculation formula is as follows:
[0085] ;
[0086] in, For the current topology graph and the candidate matching set, the first... Semantic similarity score between topological subgraphs For the current topology graph, the first The weight of each first node (determined based on the stability score of the first node; the higher the stability score, the greater the weight of the first node in semantic matching). For the current topology graph, the first The first node in the candidate matching set is the first node. The matching score of each topological subgraph This represents the number of the first node in the current topology graph. Semantic similarity score. The higher the value, the higher the value of the current topology graph and the candidate matching set. The more similar the topological subgraphs are in terms of stable semantic landmark category composition and landmark reliability, the better.
[0087] Step 6.1.4.3: For each pair of matching landmark nodes (for each pair of matching landmark nodes between the current topology graph and the current topology subgraph), compare the observation relationship edge of the first node and the observation relationship edge of the second node. If the relative second spatial distance, relative second azimuth angle, and relative second elevation angle of the first node and the second node are within the corresponding second preset range (the second preset range corresponding to the relative second spatial distance of the first node and the second node is within 1m, the second preset range corresponding to the relative second azimuth angle is within 15°, and the second preset range corresponding to the relative second elevation angle is within 15°), then the current pair of matching landmark nodes is called a geometric interior point; otherwise, it is called a geometric exterior point. The relative second azimuth angle and the relative second elevation angle are calculated using the three-dimensional centroid coordinates of the stable targets corresponding to the first node and the second node, respectively.
[0088] Step 6.1.4.4: Calculate the proportion of geometric interior points in all matching landmark node pairs, and use this proportion as the geometric similarity score between the current topological graph and the current topological subgraph. The higher the proportion of geometric interior points, the more consistent the spatial layout of the current topological graph and the current topological subgraph; the lower the proportion of geometric interior points, the more consistent the two may be in terms of semantic landmark composition, but their actual spatial positional relationship is inconsistent.
[0089] Step 6.1.5: The semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph are weighted and summed to obtain the comprehensive scene similarity between the current topological graph and the current topological subgraph. The corresponding calculation formula is as follows:
[0090] ;
[0091] in, To consider the overall scene similarity, For the current topology graph and the candidate matching set, the first... Geometric similarity score between topological subgraphs and For respectively and The weighting coefficients.
[0092] Step 6.2: For each topology graph, when comparing the maximum comprehensive scene similarity with the preset similarity value, if it is greater than the preset similarity value, then the target scene corresponding to the current topology graph is the known scene, and the topology subgraph corresponding to the maximum comprehensive scene similarity is the same scene as the target scene of the current topology graph.
[0093] If the similarity is less than the preset value, the target scene corresponding to the current topology graph is designated as an unknown scene, and the current topology graph is stored as a new topology subgraph in the historical database (the scene nodes, landmark nodes, and the relationship edges between the nodes corresponding to the current topology graph are stored in the historical database. Therefore, the construction of the historical database is to store the topology graphs corresponding to the unknown scenes into the historical database in sequence, and the historical database is continuously expanded), thereby realizing incremental graph construction of the historical database. Each topology subgraph in the historical database includes a scene node (one), landmark nodes (multiple), and the relationship edges between the nodes (where the relationship edges include the observation relationship edges between the scene node and each landmark node and the spatial relationship edges between each landmark node).
[0094] In another embodiment, to verify the effectiveness of this method, a comparative experiment was conducted with two existing methods using a public dataset (such as the SemanticKITTI dataset). The public dataset was constructed based on data collected from real road scenes and included continuous 3D point cloud data and pose trajectory information.
[0095] In this experiment, sequences 00, 02, 05, 06, 07, and 08 from the SemanticKITTI dataset were selected as experimental sequences. Each sequence corresponds to a 3D point cloud sequence continuously acquired by a mobile platform in a real road environment, containing continuous 3D point cloud data and pose trajectory information. For each sequence, keyframes were sampled in chronological order, and a portion of the keyframes were used as a historical database to store the topological subgraph of the historical database; the remaining keyframes were used as query frames to simulate the scene recognition process when the UAV re-enters a known area or passes through a similar area.
[0096] Two existing methods are Method A and Method B, and this experiment uses the scene recognition accuracy ( ), recall rate ( ) and a comprehensive evaluation of precision and recall ( This serves as an indicator for evaluating the performance of our method compared to existing methods.
[0097] Method A: This method extracts only the global appearance descriptor of the current scene node and calculates its similarity with the global appearance descriptors of each historical scene node in the historical database. The historical scene with the highest score is selected as the scene recognition result. This method is used to verify the baseline performance when scene recognition is performed solely based on global appearance information.
[0098] Method B: This method first uses a global appearance descriptor to retrieve candidate historical scenes from the historical database. Then, it further compares the semantic landmark categories, stability scores, and detection confidence scores of the current scene and the candidate historical scenes, calculates the semantic similarity score, and reorders or filters the candidate results based on the semantic similarity. This method is used to verify the effect of semantic landmark matching on improving the accuracy of scene recognition.
[0099] The comparison results are shown in Table 1:
[0100] Table 1
[0101]
[0102] As shown in Table 1, Method A only relies on global appearance descriptors for scene recognition. It is prone to mismatches in scenes with similar road structures, similar building appearances, or changes in viewing angle. Therefore, its accuracy, recall, and overall evaluation of precision and recall are relatively low.
[0103] Method B, building upon Method A, introduces semantic landmark matching. By comparing the semantic landmark composition and reliability between the current scene and candidate historical scenes, it can filter out some candidate scenes that are only visually similar but have inconsistent semantic compositions. Therefore, compared to Method A, Method B improves precision, recall, and the overall evaluation of precision and recall. However, since it has not further verified the three-dimensional spatial layout relationship between landmarks, it may still retain some pseudo-matching results that are semantically similar but have different actual spatial structures.
[0104] This method further introduces geometric verification on top of semantic matching. By comparing the relative spatial distance, relative azimuth, and relative pitch angle between matched landmark nodes, it determines the consistency of the spatial layout between the current scene and candidate historical scenes. Experimental results show that the accuracy, recall, and overall precision and recall of this method are all higher than those of methods A and B, indicating that this method can effectively eliminate false matching candidates and improve the accuracy and robustness of scene recognition through the joint constraint of semantic similarity scoring and geometric similarity scoring.
[0105] In summary, the experimental results show that, compared with methods that only use global descriptor retrieval and methods that only introduce semantic matching, the proposed method can achieve higher scene recognition accuracy, recall, and a comprehensive evaluation of precision and recall, thereby improving the reliability of scene recognition in complex scenarios.
[0106] This method selects keyframes and extracts global appearance descriptors and stable targets from them. It constructs a topology graph for each keyframe by using scene nodes, landmark nodes, and the relationships between nodes. This reduces the need for storing complete image frames and 3D point cloud data while preserving scene appearance, target semantics, and spatial information. Furthermore, by combining the topology graph with matching, it can reduce the risk of misidentification caused by single feature matching and improve the accuracy and robustness of UAV scene recognition. For the topology graph of unknown scenes that are not matched, it is stored as a new topology subgraph in the historical database, realizing incremental mapping of the historical database.
[0107] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0109] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A method for scene recognition and incremental mapping based on UAV-perceived scene topology maps, characterized in that: Acquire image frames, 3D point cloud data, and UAV pose data continuously collected by the UAV within the current preset time period for the same target scene; Filter consecutive image frames and select those that meet preset conditions as keyframes; Scene extraction and target detection are performed on each keyframe to obtain the global appearance descriptor, target region, semantic category and confidence of each keyframe, and the three-dimensional centroid coordinates of the target are calculated using the three-dimensional point cloud corresponding to the target region. The stability score of a target is calculated using the target's confidence level, and targets with a stability score that reaches the first threshold are considered stable targets, with each stable target serving as a landmark node. Construct a topology graph for each keyframe, which includes scene nodes, landmark nodes, and the relationship edges between nodes. The scene nodes contain the global appearance descriptor of the current keyframe and the corresponding UAV pose data, while the landmark nodes contain the semantic category, confidence, 3D centroid coordinates, and stability score of the corresponding stable target. Each topology graph is matched with the topology subgraphs in the historical database. The comprehensive scene similarity between each topology graph and the topology subgraphs in the historical database is calculated. For each topology graph, the maximum comprehensive scene similarity is compared with the preset similarity value. If it is greater than the preset similarity value, the target scene corresponding to the current topology graph is a known scene; otherwise, it is an unknown scene. The current topology graph is stored as a new topology subgraph in the historical database to realize incremental graph building on the historical database. The step of matching each topology graph with the topology subgraphs in the historical database and calculating the comprehensive scene similarity between each topology graph and the topology subgraphs in the historical database includes: The similarity between the global appearance descriptor in the current topology graph and the global appearance descriptors of all topology subgraphs in the historical database is calculated, and the top K topology subgraphs with the highest similarity are selected from the historical database to form a candidate matching set. Each landmark node in the current topology graph is called the first node, and each landmark node in each topology subgraph in the candidate matching set is called the second node. For each topological subgraph in the candidate matching set, each first node is matched with all second nodes in the current topological subgraph. When the semantic categories of the first node and the second node are the same and their stability scores differ within a first preset range, the first node and the second node are called a matching landmark node pair. Based on all matching landmark node pairs between the current topological graph and the current topological subgraph, calculate the semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph; The semantic similarity score and geometric similarity score between the current topological graph and the current topological subgraph are weighted and summed to obtain the comprehensive scene similarity between the current topological graph and the current topological subgraph. The step of calculating semantic similarity scores and geometric similarity scores between the current topological graph and the current topological subgraph based on all matching landmark node pairs between the current topological graph and the current topological subgraph includes: Calculate the matching score of the first node in each matching landmark node pair, and the matching score is equal to the weighted sum of the confidence and stability scores of the first node in the current matching landmark node pair; The semantic similarity score between the current topological graph and the current topological subgraph is obtained by weighted averaging of the matching scores of all first nodes in the current topological graph. For each matching landmark node pair, compare the observation relationship edge of the first node and the observation relationship edge of the second node. If the relative second spatial distance, relative second azimuth angle and relative second elevation angle of the first node and the second node are respectively within the corresponding second preset range, then the current matching landmark node pair is called the geometric interior point; otherwise, it is called the geometric exterior point. The relative second azimuth angle and relative second elevation angle are calculated by the three-dimensional centroid coordinates of the stable targets corresponding to the first node and the second node, respectively. Calculate the proportion of geometric interior points in all matching landmark node pairs, and use this proportion as the geometric similarity score between the current topology graph and the current topology subgraph.
2. The scene recognition and incremental mapping method based on UAV-perceived scene topology map as described in claim 1, characterized in that: The UAV pose data includes the UAV's position and attitude, and the attitude includes the heading angle and pitch angle.
3. The scene recognition and incremental mapping method based on UAV-perceived scene topology map as described in claim 2, characterized in that: The step of filtering consecutive image frames and selecting image frames that meet preset conditions as keyframes includes: Step 1: Use the first image frame in a series of consecutive image frames as the initial keyframe; Step II: Each image frame following the current keyframe is called a candidate image frame, and the first score between each candidate image frame and the current keyframe is calculated in turn. If the first score is greater than the second threshold, the current candidate image frame is used as the new current keyframe; otherwise, the current candidate image frame is not used as a keyframe. Step III: Repeat Step II and iterate continuously until all image frames in a series of consecutive image frames have participated in the calculation of the first score at least once. Then stop the iteration and complete the screening of key frames in the series of consecutive image frames. The first score is the weighted sum of the changes in UAV position, UAV attitude, and time between the candidate image frame and the current keyframe.
4. The scene recognition and incremental mapping method based on UAV-perceived scene topology map as described in claim 1, characterized in that: When performing scene extraction and target detection on each keyframe, the image processing model is used to process the keyframe to obtain the global appearance descriptor of the current keyframe, and the target detection model is used to process the keyframe to obtain the target region, semantic category and confidence of each target in the current keyframe. When using the three-dimensional point cloud data corresponding to the target area to calculate the three-dimensional centroid coordinates of the target, the three-dimensional point cloud data is projected into the image frame. The three-dimensional point cloud falling into the target area is taken as the point cloud of the target, and the average coordinate of all the point clouds of the target is taken as the three-dimensional centroid coordinates of the target.
5. The scene recognition and incremental mapping method based on UAV-perceived scene topology map as described in claim 1, characterized in that: When calculating the stability score of a target, the target's confidence level is weighted and summed with the number of times the target is observed in adjacent frames to obtain the stability score. The number of times the target is observed in adjacent frames is the sum of the number of times the target appears in the current keyframe, the previous keyframe, and the next keyframe.
6. The scene recognition and incremental mapping method based on UAV-perceived scene topology map as described in claim 2, characterized in that: The relationship edges of each node in the topology graph of each keyframe include the observation relationship edges between scene nodes and each landmark node, as well as the spatial relationship edges between each landmark node; The observation relationship edge includes the relative first spatial distance, relative first azimuth angle, and relative first pitch angle between the UAV in the scene node and the stable target corresponding to the landmark node. The relative first spatial distance, relative first azimuth angle, and relative first pitch angle are all calculated by the position of the UAV in the scene node and the three-dimensional centroid coordinates of the stable target corresponding to the landmark node. The spatial relationship edge includes the relative second spatial distance between two landmark nodes, and the relative second spatial distance is calculated using the three-dimensional centroid coordinates of the stable targets corresponding to the two landmark nodes.
7. The scene recognition and incremental mapping method based on UAV-perceived scene topology map as described in claim 1, characterized in that: Each topological subgraph in the historical database includes scene nodes, landmark nodes, and the relationship edges between the nodes.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative mapping and sensing method and system based on semantic consistency
CN117152249A
System and method of hybrid scene representation for visual simultaneous localization and mapping
US20240104771A1