Indoor positioning method and system for airport intelligent scooter based on ground grid vision
Patent Information
- Application Number
- CN202610819328.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]本发明的目的就在于解决仅依靠导航检测,使得定位精度不足;难以获得精确的位置信息;且机场地面纹理复杂和光照不均,导致地面特征提取不足,导致定位的精准度偏低的问题,而提出基于地面网格视觉的机场智能代步车室内定位方法及系统
Smart Images

Figure CN122729993A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of indoor positioning technology, specifically relating to an indoor positioning method and system for airport intelligent mobility vehicles based on ground grid vision. Background Technology
[0002] With the continuous growth of airport passenger traffic, the demand for indoor mobility aids is becoming increasingly urgent for elderly passengers, people with mobility impairments, and passengers carrying large luggage. The existing intelligent mobility vehicles mainly use wireless signals, visual SLAM, and lidar positioning technologies to support the autonomous positioning and navigation of airport intelligent mobility vehicles.
[0003] Existing technology (publication number: CN116222543A) provides a multi-sensor fusion map construction method and system for robot environmental perception. It addresses the traditional method of distributing data across different SLAM algorithms. By acquiring left and right eye image information through a binocular camera, binocular vision technology can estimate the robot's pose and scene depth information to construct a map. Simultaneously, high-precision depth information is obtained using LiDAR to optimize the map, improving its accuracy and robustness. Secondly, IMU measurement information can be used for robot motion estimation and error compensation, improving the accuracy and robustness of quadruped robot motion. Finally, by fusing the measurement information from the binocular camera, LiDAR, and IMU sensor, more accurate and robust localization and mapping of the quadruped robot are achieved.
[0004] The aforementioned patent achieves the localization of a quadruped robot. However, in the actual localization process, the positioning accuracy obtained by relying solely on inertial navigation detection is insufficient, making it difficult to obtain accurate location information. Furthermore, the complex texture of the airport ground and uneven lighting result in insufficient extraction of ground features, leading to low positioning accuracy. Summary of the Invention
[0005] The purpose of this invention is to solve the problems that relying solely on navigation detection results in insufficient positioning accuracy; difficulty in obtaining accurate location information; and the complex texture of airport ground and uneven lighting lead to insufficient ground feature extraction, resulting in low positioning accuracy. Therefore, this invention proposes an indoor positioning method and system for airport intelligent mobility vehicles based on ground grid vision.
[0006] In a first aspect of this invention, an indoor positioning method for intelligent mobility vehicles at airports based on ground-grid vision is first proposed, the method comprising: During the operation of the mobility scooter, inertial navigation data and continuous frames of ground images are acquired; After preprocessing the target ground image, a candidate edge map is obtained by processing it according to the adaptive gray-scale gradient analysis algorithm; after performing grid line fitting processing on the candidate edge map, the set of intersection points of the grid lines is calculated; after performing pixel coordinate transformation on the set of intersection points of the grid lines, a ground coordinate set is obtained; the target ground image is any ground image in a series of ground images; The theoretical positioning result of the mobility scooter is predicted based on its initial position and inertial navigation data. The theoretical localization results are used to determine the set of candidate matching nodes corresponding to the theoretical ground image in the indoor semantic feature map; The visual positioning result of the mobility scooter is obtained by matching and correcting the candidate matching node set based on the ground coordinate set.
[0007] Optionally, before the mobility scooter is put into operation, the construction of the indoor semantic feature map includes the following steps: Acquire ground point cloud data and indoor ground image data, and perform ground information enhancement on the indoor ground image data using the ground point cloud data to obtain enhanced image data; and establish a global ground coordinate system with a preset origin as the reference point, the x-axis pointing east and the y-axis pointing north. Based on the enhanced image data, the indoor floor tile grid features are labeled, and all grid lines formed by tile joints and their intersections are extracted as grid nodes; the two-dimensional coordinates of each grid node in the global coordinate system are recorded to construct a grid node coordinate database. The semantic region annotation data is obtained by semantically annotating the indoor ground based on the indoor ground image data. The specific process of semantic region annotation is as follows: the indoor ground is divided into multiple semantic regions, and the data of each semantic region includes region type, access rules and channel width. Simultaneously, the location and type information of fixed ground features are labeled as auxiliary positioning features; The grid node coordinate database, semantic region annotation data, and auxiliary feature data are merged to construct a complete indoor semantic feature map.
[0008] Optionally, after preprocessing the target ground image, the candidate edge map can be obtained by processing it using an adaptive gray-level gradient analysis algorithm, including: The target ground image is converted to grayscale to obtain a grayscale image; the grayscale image is normalized to obtain a first preprocessed image; the first preprocessed image is denoised by Gaussian filtering to obtain a second preprocessed image. The second preprocessed image is adaptively edge-trimmed to obtain a binary image; After constrained filtering of the binary image, each pixel in the horizontal and vertical gradient directions is retained; non-maximum suppression is then applied to each pixel in the horizontal and vertical gradient directions to generate a candidate edge map.
[0009] Optionally, the adaptive edgeification steps include: The gradient operator is used to calculate the horizontal and vertical gradient values of each pixel in the second preprocessed image, and the gradient magnitude and gradient direction of each pixel are calculated based on the horizontal and vertical gradient values of each pixel. The second preprocessed image is divided into multiple sub-regions. The mean and standard deviation of the gradient magnitude of each sub-region are calculated. The seam edge value of the corresponding sub-region is calculated based on the mean and standard deviation of the gradient magnitude. If the seam edge value is greater than the preset threshold, then each pixel in the corresponding sub-region is recorded as an edge candidate point. If the seam edge value is less than the preset threshold, then each pixel in the corresponding sub-region is recorded as a non-edge candidate point. The binary image is obtained by combining all edge candidate points and all non-edge candidate points.
[0010] Optionally, the step of performing grid line fitting processing on the candidate edge map and then calculating the set of grid line intersections includes: A set of candidate lines is obtained by performing Hough transform line detection on the candidate edge map; Cluster the candidate line set to obtain the refined line set; divide the refined line set into a horizontal line subset and a vertical line subset; After sorting the lines in each subset according to their positions, calculate the coordinates of all intersection points of the horizontal and vertical line subsets to obtain the set of intersection points of the grid lines in the current frame image.
[0011] Optionally, the visual positioning results of the mobility scooter can be obtained by matching and correcting the candidate matching node set based on the ground coordinate set, including: Calculate the distance and azimuth between any two adjacent ground coordinate points in the ground coordinate set and record them in the first structure label; construct a ground topology map based on all ground coordinate points and their corresponding first structure labels; Calculate the distance and azimuth between any two adjacent nodes in the candidate matching node set and record them in the second structure label; construct a candidate topology graph based on all candidate matching nodes and their corresponding second structure labels; The matching results are obtained by performing structural matching based on the ground topology map and the candidate topology map; The theoretical positioning result of the mobility scooter is corrected based on the matching result to obtain the visual positioning result.
[0012] Optionally, the step of obtaining the matching result by performing structure matching based on the ground topology map and the candidate topology map includes: Query all boundary nodes in the ground topology map that have adjacent grid nodes in only one of the four adjacent horizontal or vertical directions; connect all boundary nodes to obtain the boundary polygon, and calculate the geometric features of the boundary polygon; the geometric features include length features, angle features, and total perimeter features of the boundary. Traverse all nodes in the candidate topology graph, find all candidate polygons whose perimeter is less than the difference between the perimeter and the total perimeter of the boundary and the preset perimeter threshold, and record them in the candidate window. The length matching value is obtained by calculating the similarity between the length feature of the boundary polygon and the length feature of each perimeter polygon in the candidate window. If the length matching value is within the preset length matching range, the current perimeter polygon in the candidate window is selected; all nodes in the boundary polygon are matched with all nodes in the current perimeter polygon to obtain the matching result; the matching result is all points in the boundary polygon that have completed the matching.
[0013] Optionally, the step of correcting the theoretical positioning result of the mobility scooter based on the matching result to obtain the visual positioning result includes: Based on the matching results, determine the corresponding node of each point in the ground coordinate set in the candidate matching node set, calculate the coordinate offset and angle offset between all corresponding point pairs, and add the coordinate offset and angle offset to the theoretical positioning result to obtain the visual positioning result of the mobility scooter.
[0014] In a second aspect of this invention, an indoor positioning system for intelligent mobility scooters at airports based on ground-grid vision is proposed, the system comprising: Data acquisition module: Acquires inertial navigation data and continuous frames of ground images while the mobility scooter is in operation; Ground coordinate module: After preprocessing the target ground image, candidate edge maps are obtained by processing it according to the adaptive gray-scale gradient analysis algorithm; after grid line fitting processing of the candidate edge maps, the set of intersection points of the grid lines is calculated; after pixel coordinate transformation of the set of intersection points of the grid lines, the ground coordinate set is obtained; the target ground image is any ground image in the continuous frame ground images; Inertial navigation module: predicts the theoretical positioning result of the mobility scooter based on its initial position and inertial navigation data; Visual positioning module: Determines the set of candidate matching nodes corresponding to the theoretical ground image in the indoor semantic feature map based on the theoretical positioning results; and performs matching correction on the set of candidate matching nodes based on the set of ground coordinates to obtain the visual positioning results of the mobility scooter.
[0015] The beneficial effects of this invention are: This invention proposes an indoor positioning method and system for intelligent mobility scooters at airports based on ground grid vision. Through an adaptive grayscale gradient analysis algorithm combined with grid line fitting, it can robustly extract key points of grid intersections formed by tile seams in airport ground images with uneven lighting distribution and complex texture features. This effectively avoids external interference such as ground reflection and texture homogenization, significantly enhancing the stability and accuracy of feature extraction. By relying on the estimated positioning information output by inertial navigation to reduce the map matching retrieval interval, and then by constructing a topological structure of ground feature coordinate points and performing topological structure matching with candidate map nodes, high-precision correction of inertial navigation positioning deviations is achieved, avoiding the positioning jitter problem that easily arises from relying solely on visual feature matching. This achieves low-cost, high-precision indoor positioning at airports, ensuring continuous and uninterrupted positioning for the mobility scooter during long-distance travel. Attached Figure Description
[0016] The invention will now be further described with reference to the accompanying drawings.
[0017] Figure 1 A flowchart of an indoor positioning method for airport intelligent mobility vehicles based on ground grid vision provided in an embodiment of the present invention; Figure 2 This is a framework diagram of an indoor positioning system for airport intelligent mobility vehicles based on ground grid vision, provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and B can represent: A alone, A and B simultaneously, and B alone. Furthermore, descriptions involving "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" can explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0019] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] This invention provides an indoor positioning method for intelligent mobility vehicles at airports based on ground-grid vision. See also... Figure 1 , Figure 1 A flowchart illustrating an indoor positioning method for intelligent airport mobility vehicles based on ground-grid vision, provided in an embodiment of the present invention. The method includes the following steps: During the operation of the mobility scooter, inertial navigation data and continuous frames of ground images are acquired; After preprocessing the target ground image, a candidate edge map is obtained by processing it according to the adaptive gray-level gradient analysis algorithm; after mesh line fitting processing of the candidate edge map, the set of intersection points of the mesh lines is calculated; after pixel coordinate transformation of the set of intersection points of the mesh lines, the ground coordinate set is obtained; the target ground image is any ground image in the continuous frame ground images; The theoretical positioning result of the mobility scooter is predicted based on its initial position and inertial navigation data. The theoretical localization results are used to determine the set of candidate matching nodes corresponding to the theoretical ground image in the indoor semantic feature map; The visual positioning result of the mobility scooter is obtained by matching and correcting the candidate matching node set based on the ground coordinate set.
[0021] The indoor positioning method for airport smart mobility scooters based on ground grid vision provided in this invention utilizes an adaptive grayscale gradient analysis algorithm and grid line fitting to stably extract grid line intersections formed by tile seams from airport ground images with uneven lighting and complex textures. This effectively overcomes interference factors such as ground reflection and monotonous texture, significantly improving the robustness and accuracy of feature extraction. The theoretical positioning results provided by inertial navigation are used to narrow the search range for map matching. Then, by constructing a topological graph of the ground coordinate point set and performing structural matching with candidate nodes in the map, accurate correction of inertial navigation is achieved, avoiding positioning fluctuations caused by relying solely on visual feature matching. This achieves high-precision, low-cost indoor positioning at airports, ensuring the continuity of positioning for the mobility scooter during long-distance operation.
[0022] In one implementation, the construction of an indoor semantic feature map before the mobility scooter is put into operation includes the following steps: Acquire ground point cloud data and indoor ground image data (specifically, images of the airport ground). Use the ground point cloud data to perform ground information enhancement on the indoor ground image data to obtain enhanced image data. Establish a global ground coordinate system with a preset origin as the reference point, x-axis pointing east and y-axis pointing north. Based on the enhanced image data, the indoor floor tile grid features are labeled, and all grid lines formed by tile joints and their intersections are extracted as grid nodes; the two-dimensional coordinates of each grid node in the global coordinate system are recorded to construct a grid node coordinate database. The semantic region annotation data is obtained by semantically annotating the indoor ground based on the indoor ground image data. The specific process of semantic region annotation is as follows: the indoor ground is divided into multiple semantic regions, and the data of each semantic region includes region type, passage rules and channel width. Simultaneously, the location and type information of fixed ground features are labeled as auxiliary positioning features; By merging the grid node coordinate database, semantic region annotation data, and auxiliary feature data, a complete indoor semantic feature map is constructed.
[0023] In one implementation, point cloud data provides realistic 3D spatial information, which can be used to reverse-correct geometric distortions in images and enhance the geometric accuracy of ground seams. Simultaneously, dividing the map into three layers—a grid node database, semantic regions (such as boarding gates, corridors, and elevator entrances), and fixed features (such as pillars and counters)—provides multi-dimensional constraints for subsequent positioning: grid nodes are used for precise location matching, semantic regions for path planning and compliance checks, and fixed features for backup corrections in case of positioning failure. The resulting map not only supports high-precision visual positioning but also guides vehicles to comply with airport traffic rules (such as no-fly zones and speed-limited areas), improving the system's safety and intelligence.
[0024] In one specific embodiment, at the departure level of an airport terminal, staff used a handheld LiDAR and an industrial camera to collect data along a planned route. After fusing the point cloud data with the image data, all intersections of the 800mm×800mm tile joints were extracted, and their global coordinates were recorded (e.g., intersection A: x=12.5m, y=30.2m). Simultaneously, the waiting area was marked as a low-speed passage zone, the sides of the moving walkway were marked as no-stopping zones, and the outlines of three load-bearing columns were stored as auxiliary features in the map. This map was subsequently used for the real-time positioning of the mobility scooters.
[0025] In one implementation, after preprocessing the target ground image, a candidate edge map is obtained by processing it using an adaptive gray-level gradient analysis algorithm, including: The target ground image is converted to grayscale to obtain a grayscale image; the grayscale image is normalized to obtain a first preprocessed image; the first preprocessed image is denoised by Gaussian filtering to obtain a second preprocessed image. The second preprocessed image is adaptively edge-trimmed to obtain a binary image; After constrained filtering of the binary image, each pixel in the horizontal and vertical gradient directions is retained; non-maximum suppression is then applied to each pixel in the horizontal and vertical gradient directions to generate a candidate edge map.
[0026] Specifically, the candidate edge map is composed of grid lines with a width of one pixel; In one implementation, grayscale conversion and illumination normalization (such as CLAHE) eliminate global illumination differences, ensuring stable contrast for tile seams across different brightness regions. Gaussian filtering suppresses random noise in image acquisition, avoiding false edge interference. Adaptive edge detection dynamically determines edge thresholds based on local gradient statistical features, avoiding over-detection in high-contrast areas and under-detection in low-contrast areas due to fixed thresholds, thus fully preserving tile seams in areas with weak texture. Constraint filtering retains only gradient pixels in the horizontal and vertical directions, eliminating oblique false edges (such as tile patterns and reflective spots), significantly reducing interference from non-mesh structures. Non-maximum suppression refines edges to a single pixel width, generating continuous and clear candidate edge maps, providing high-purity, low-redundancy input for subsequent mesh line fitting. Ultimately, this method can stably extract complete tile mesh edges even in airport environments with drastic lighting changes, ground reflection, or partial wear, significantly improving the localization accuracy of feature detection.
[0027] In one implementation, illumination normalization employs adaptive histogram equalization (CLAHE); CLAHE parameter settings: block size is 8×8, contrast limit threshold is 2.0-3.0; Gaussian filtering is used for noise reduction, with a filter kernel size of 5×5 and standard deviation σ=1.0-1.5; In one implementation, the specific steps of adaptive edge detection include: The gradient operator is used to calculate the horizontal and vertical gradient values of each pixel in the second preprocessed image, and the gradient magnitude and gradient direction of each pixel are calculated based on the horizontal and vertical gradient values of each pixel. The second preprocessed image is divided into multiple sub-regions. The mean and standard deviation of the gradient magnitude of each sub-region are calculated. The seam edge value of the corresponding sub-region is calculated based on the mean and standard deviation of the gradient magnitude. If the seam edge value is greater than the preset threshold, then each pixel in the corresponding sub-region is recorded as an edge candidate point. If the seam edge value is less than the preset threshold, then each pixel in the corresponding sub-region is recorded as a non-edge candidate point. The binary image is obtained by combining all edge candidate points and all non-edge candidate points.
[0028] In one implementation, in an airport ground image, the texture contrast of different areas varies greatly, and the joints of the paving stones in well-lit areas are clear, while the joints in shadowed areas have very low contrast. By dividing the image into sub-regions and statistically analyzing the local gradient distribution, it is possible to dynamically adjust whether each region belongs to an edge region, making edge detection robust to changes in local lighting and texture. This approach can completely preserve the joints of the paving stones in shadowed areas while suppressing pseudo-texture interference on smooth paving stone surfaces. The resulting binary image has continuous grid lines with few breaks, significantly improving the accuracy of subsequent line detection and intersection calculation.
[0029] In one implementation, after performing grid line fitting on the candidate edge map, the set of grid line intersections is calculated, including: A set of candidate lines is obtained by performing Hough transform line detection on the candidate edge map; Cluster the candidate line set to obtain the refined line set; divide the refined line set into a horizontal line subset and a vertical line subset; After sorting the lines in each subset according to their positions, calculate the coordinates of all intersection points of the horizontal and vertical line subsets to obtain the set of intersection points of the grid lines in the current frame image.
[0030] In one implementation, the Hough transform line detection parameters are set as follows: distance resolution ρ = 1 pixel, angle resolution θ = π / 180, and the cumulative threshold is adaptively set according to the image size, specifically calculated as min(image width, image height) × 0.15.
[0031] In one implementation, redundancy can be effectively removed by clustering and merging lines with similar angles and intercepts, resulting in a unique representative line for each actual seam. After sorting by position, the intersection points are calculated, ensuring that the order of the intersection points matches the actual arrangement of the paving stones. The generated set of intersection points is of moderate size (typically tens to hundreds per frame), and the adjacency relationships of the intersection points are clear, facilitating subsequent construction of the topology graph. Furthermore, the sorted intersection points directly reflect the row and column structure of the grid in the vehicle's forward direction, providing a basis for heading angle estimation.
[0032] In one implementation, the visual positioning result of the mobility scooter is obtained by matching and correcting the candidate matching node set based on the ground coordinate set, including: Calculate the distance and azimuth between any two adjacent ground coordinate points in the ground coordinate set and record them in the first structure label; construct a ground topology map based on all ground coordinate points and their corresponding first structure labels; Calculate the distance and azimuth between any two adjacent nodes in the candidate matching node set and record them in the second structure label; construct a candidate topology graph based on all candidate matching nodes and their corresponding second structure labels; The matching results are obtained by performing structural matching based on the ground topology map and the candidate topology map; The theoretical positioning result of the mobility scooter is corrected based on the matching result to obtain the visual positioning result.
[0033] In one implementation, the intersection set is abstracted into a graph structure (nodes are intersections, edges are distances and azimuths), transforming the localization problem into a graph matching problem, which fully utilizes the geometric invariance of the mesh. The advantage of this approach is that even if only some intersections are detected in the image (e.g., occluded by luggage or pedestrians), the topological graph can still find the corresponding region in the map through local structure matching; and the matching process is unaffected by global coordinate offsets and rotations.
[0034] In a specific embodiment, the mobility scooter detects 12 intersection points (3 rows, 4 columns) in the current frame. The distances between adjacent intersection points are calculated: horizontally adjacent points are both 0.6m, vertically adjacent points are both 0.6m, and azimuth angles are both 0° or 90°. These relationships are stored in the first structural label, forming a small grid topology map. Near the theoretical positioning result provided by inertial navigation, 50 candidate nodes are extracted from the semantic map to construct a candidate topology map. Structural matching reveals that the current topology map perfectly matches the side length and angle of a 3×4 sub-block in the candidate topology map, successfully confirming the matching relationship.
[0035] In one implementation, the matching result obtained by performing structure matching based on the ground topology map and the candidate topology map includes: Query all boundary nodes in the ground topology map that have adjacent grid nodes in only one of the four adjacent horizontal or vertical directions; connect all boundary nodes to obtain the boundary polygon, and calculate the geometric features of the boundary polygon; the geometric features include length features, angle features, and total perimeter of the boundary. Traverse all nodes in the candidate topology graph, find all candidate polygons whose perimeter is less than the difference between the perimeter and the total perimeter of the boundary and the preset perimeter threshold, and record them in the candidate window. The length matching value is obtained by calculating the similarity between the length feature of the boundary polygon and the length feature of each perimeter polygon in the candidate window. If the length matching value is within the preset length matching range, the current perimeter polygon in the candidate window is selected; all nodes in the boundary polygon are matched with all nodes in the current perimeter polygon to obtain the matching result; the matching result is all points in the boundary polygon that have completed the matching.
[0036] Specifically, the similarity process is as follows: if the difference between the perimeter of the candidate topology graph and the total perimeter of the boundary is within a preset difference threshold, then it is considered similar; the candidate window specifically includes the length sequence, angle sequence, and all nodes inside the perimeter of the selected boundary; the similarity calculation uses methods such as cosine similarity.
[0037] In one implementation, by first coarsely screening regions with similar perimeters and then finely comparing the side length sequences, the matching search range can be greatly reduced. This reduces the complexity of quadratic matching from O(N) to O(N). ^2The number of candidate windows is reduced to O(K·L) (K is the number of candidate windows and L is the number of boundary nodes), which significantly improves real-time performance; since it only matches the boundaries, it is not sensitive to missing internal nodes (such as those that are occluded), which improves the matching success rate.
[0038] In a specific embodiment, the grid detected in the current frame is an "L"-shaped boundary (missing corner region). The boundary polygon has 8 nodes, a total perimeter of 12.8m, and a side length sequence of [0.6, 0.6, 0.6, 1.2, 0.6, 0.6, 0.6, 1.2]. All node combinations are traversed in the candidate topology graph to find 3 polygons with perimeters between 12.6 and 13.0 meters. The similarity of the side length sequences is then compared, and one with a similarity of 95% is identified as the matching region.
[0039] In one implementation, the visual positioning result is obtained by correcting the theoretical positioning result of the mobility scooter based on the matching result, including: Based on the matching results, determine the corresponding node of each point in the ground coordinate set in the candidate matching node set, calculate the coordinate offset and angle offset between all corresponding point pairs, and add the coordinate offset and angle offset to the theoretical positioning result to obtain the visual positioning result of the mobility scooter.
[0040] In one implementation, based on the matching results, the set of ground coordinates P = {p1, p2, ..., p...} detected from the current frame image is... m The set of candidate matching nodes Q={q1,q2,...,q} in the indoor semantic feature map. n A one-to-one correspondence is established between}. The matching results are usually derived from topological graph structure matching (such as boundary polygon matching or node adjacency matching), and the final output is a list of corresponding point pairs {(p i ,q i )}, where p i It is the intersection point of the image detection, q i This refers to the actual coordinate nodes of the vehicle on the map. For all matching point pairs, the coordinate difference of each pair in the global coordinate system is calculated. Then, the average of all coordinate differences is taken to obtain the final horizontal and vertical offsets. For each matching point pair, the angle difference between the theoretically inferred heading of the vehicle from inertial navigation and the direction of the grid lines on the map is calculated. The angle deviation is calculated using the vector formed by any two points in the matching point pair; specifically, two matching point pairs can be selected, and the included angle between the two matching point pairs can be calculated. The average of all included angles is taken to obtain the angle offset. The final horizontal offset, final vertical offset, and angle offset are integrated to obtain the corrected positioning result. The theoretical positioning result and the corrected positioning result are added together to obtain the visual positioning result of the vehicle.
[0041] Based on the same inventive concept, embodiments of the present invention also provide an indoor positioning system for intelligent airport mobility vehicles based on ground grid vision. See also Figure 2 , Figure 2 A framework diagram of an airport intelligent mobility scooter indoor positioning system based on ground grid vision provided for embodiments of the present invention includes: Data acquisition module: Acquires inertial navigation data and continuous frames of ground images while the mobility scooter is in operation; Ground coordinate module: After preprocessing the target ground image, it is processed by an adaptive gray-scale gradient analysis algorithm to obtain a candidate edge map; after performing grid line fitting on the candidate edge map, the set of grid line intersection points is calculated; after performing pixel coordinate transformation on the set of grid line intersection points, the ground coordinate set is obtained; the target ground image is any ground image in a series of ground images; Inertial navigation module: predicts the theoretical positioning result of the mobility scooter based on its initial position and inertial navigation data; Visual positioning module: Determines the set of candidate matching nodes corresponding to the theoretical ground image in the indoor semantic feature map based on the theoretical positioning results; and performs matching correction on the set of candidate matching nodes based on the set of ground coordinates to obtain the visual positioning results of the mobility scooter.
[0042] The airport intelligent mobility scooter indoor positioning system based on ground grid vision provided in this invention utilizes an adaptive grayscale gradient analysis algorithm and grid line fitting processing to stably extract grid line intersections formed by tile seams from airport ground images with uneven lighting and complex textures. This effectively overcomes interference factors such as ground reflection and monotonous texture, significantly improving the robustness and accuracy of feature extraction. The theoretical positioning results provided by inertial navigation are used to narrow down the search range for map matching. Then, by constructing a topological map of the ground coordinate point set and structurally matching it with candidate nodes in the map, accurate correction of inertial navigation is achieved, avoiding positioning fluctuations caused by relying solely on visual feature matching. This system achieves high-precision, low-cost airport indoor positioning, ensuring the continuity of positioning for the mobility scooter during long-distance operation.
[0043] The foregoing has described one embodiment of the present invention in detail, but this content is merely a preferred embodiment and should not be considered as limiting the scope of the present invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the scope of the claims of this invention.
Claims
1. An indoor positioning method for intelligent mobility vehicles at airports based on ground-grid vision, characterized in that, The method includes: During the operation of the mobility scooter, inertial navigation data and continuous frames of ground images are acquired; After preprocessing the target ground image, a candidate edge map is obtained by processing it according to the adaptive gray-scale gradient analysis algorithm; after performing grid line fitting processing on the candidate edge map, the set of intersection points of the grid lines is calculated; after performing pixel coordinate transformation on the set of intersection points of the grid lines, a ground coordinate set is obtained; the target ground image is any ground image in a series of ground images; The theoretical positioning result of the mobility scooter is predicted based on its initial position and inertial navigation data. The theoretical localization results are used to determine the set of candidate matching nodes corresponding to the theoretical ground image in the indoor semantic feature map; The visual positioning result of the mobility scooter is obtained by matching and correcting the candidate matching node set based on the ground coordinate set.
2. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 1, characterized in that, Before the mobility scooter is put into operation, the construction of the indoor semantic feature map includes the following specific steps: Acquire ground point cloud data and indoor ground image data, and perform ground information enhancement on the indoor ground image data using the ground point cloud data to obtain enhanced image data; and establish a global ground coordinate system with a preset origin as the reference point, the x-axis pointing east and the y-axis pointing north. Based on the enhanced image data, the indoor floor tile grid features are labeled, and all grid lines formed by tile joints and their intersections are extracted as grid nodes; the two-dimensional coordinates of each grid node in the global coordinate system are recorded to construct a grid node coordinate database. The semantic region annotation data is obtained by semantically annotating the indoor ground based on the indoor ground image data. The specific process of semantic region annotation is as follows: the indoor ground is divided into multiple semantic regions, and the data of each semantic region includes region type, access rules and channel width. Simultaneously, the location and type information of fixed ground features are labeled as auxiliary positioning features; The grid node coordinate database, semantic region annotation data, and auxiliary feature data are merged to construct a complete indoor semantic feature map.
3. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 1, characterized in that, After preprocessing the target ground image, the candidate edge map is obtained by further processing it using an adaptive gray-scale gradient analysis algorithm, including: The target ground image is converted to grayscale to obtain a grayscale image; the grayscale image is normalized to obtain a first preprocessed image; the first preprocessed image is denoised by Gaussian filtering to obtain a second preprocessed image. The second preprocessed image is adaptively edge-trimmed to obtain a binary image; After constrained filtering of the binary image, each pixel in the horizontal and vertical gradient directions is retained; non-maximum suppression is then applied to each pixel in the horizontal and vertical gradient directions to generate a candidate edge map.
4. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 3, characterized in that, The specific steps of the adaptive edgeification include: The gradient operator calculates the horizontal and vertical gradient values of each pixel in the second preprocessed image, and calculates the gradient magnitude and gradient direction of each pixel based on the horizontal and vertical gradient values of each pixel. The second preprocessed image is divided into multiple sub-regions. The mean and standard deviation of the gradient magnitude of each sub-region are calculated. The seam edge value of the corresponding sub-region is calculated based on the mean and standard deviation of the gradient magnitude. If the seam edge value is greater than the preset threshold, then each pixel in the corresponding sub-region is recorded as an edge candidate point; If the seam edge value is less than the preset threshold, then each pixel in the corresponding sub-region is recorded as a non-edge candidate point. The binary image is obtained by combining all edge candidate points and all non-edge candidate points.
5. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 1, characterized in that, The process of performing grid line fitting on the candidate edge map and then calculating the set of grid line intersections includes: A set of candidate lines is obtained by performing Hough transform line detection on the candidate edge map; Cluster the candidate line set to obtain the refined line set; divide the refined line set into a horizontal line subset and a vertical line subset; After sorting the lines in each subset according to their positions, calculate the coordinates of all intersection points of the horizontal and vertical line subsets to obtain the set of intersection points of the grid lines in the current frame image.
6. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 1, characterized in that, The visual positioning results of the mobility scooter are obtained by matching and correcting the candidate matching node set based on the ground coordinate set, including: Calculate the distance and azimuth between any two adjacent ground coordinate points in the ground coordinate set and record them in the first structure label; construct a ground topology map based on all ground coordinate points and their corresponding first structure labels; Calculate the distance and azimuth between any two adjacent nodes in the candidate matching node set and record them in the second structure label; construct a candidate topology graph based on all candidate matching nodes and their corresponding second structure labels; The matching results are obtained by performing structural matching based on the ground topology map and the candidate topology map; The theoretical positioning result of the mobility scooter is corrected based on the matching result to obtain the visual positioning result.
7. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 6, characterized in that, The process of obtaining matching results by performing structural matching based on the ground topology map and candidate topology maps includes: Query all boundary nodes in the ground topology map that have adjacent grid nodes in only one of the four adjacent horizontal or vertical directions; connect all boundary nodes to obtain the boundary polygon, and calculate the geometric features of the boundary polygon; the geometric features include length features, angle features, and total perimeter features of the boundary. Traverse all nodes in the candidate topology graph, find all candidate polygons whose perimeter is less than the difference between the perimeter and the total perimeter of the boundary and the preset perimeter threshold, and record them in the candidate window. The length matching value is obtained by calculating the similarity between the length feature of the boundary polygon and the length feature of each perimeter polygon in the candidate window. If the length matching value is within the preset length matching range, the current perimeter polygon in the candidate window is selected; all nodes in the boundary polygon are matched with all nodes in the current perimeter polygon to obtain the matching result; the matching result is all points in the boundary polygon that have completed the matching.
8. The indoor positioning method for airport intelligent mobility vehicles based on ground grid vision according to claim 6, characterized in that, The process of correcting the theoretical positioning result of the mobility scooter based on the matching result to obtain the visual positioning result includes: Based on the matching results, determine the corresponding node of each point in the ground coordinate set in the candidate matching node set, calculate the coordinate offset and angle offset between all corresponding point pairs, and add the coordinate offset and angle offset to the theoretical positioning result to obtain the visual positioning result of the mobility scooter.
9. An indoor positioning system for airport intelligent mobility scooters based on terrestrial grid vision, used to implement the indoor positioning method for airport intelligent mobility scooters based on terrestrial grid vision as described in any one of claims 1-8, characterized in that, The system includes: Data acquisition module: Acquires inertial navigation data and continuous frames of ground images while the mobility scooter is in operation; Ground coordinate module: After preprocessing the target ground image, candidate edge maps are obtained by processing it according to the adaptive gray-scale gradient analysis algorithm; after grid line fitting processing of the candidate edge maps, the set of intersection points of the grid lines is calculated; after pixel coordinate transformation of the set of intersection points of the grid lines, the ground coordinate set is obtained; the target ground image is any ground image in the continuous frame ground images; Inertial navigation module: predicts the theoretical positioning result of the vehicle based on its initial position and inertial navigation data; Visual positioning module: determines the candidate matching node set corresponding to the theoretical ground image in the indoor semantic feature map based on the theoretical positioning result; and performs matching correction on the candidate matching node set based on the ground coordinate set to obtain the visual positioning result of the vehicle.
Citation Information
Patent Citations
Multi-sensor fusion map construction method and system for robot environment perception
CN116222543A