Micro unmanned aerial vehicle indoor positioning method and system based on ultrasonic environment perception

By combining multi-source data fusion localization methods from visual cameras, IMUs, and ToF sensors, and identifying and correcting repetitive texture scenes, the problem of error accumulation in indoor positioning of micro UAVs is solved, and stable indoor positioning and obstacle avoidance functions are achieved.

CN121498718AActive Publication Date: 2026-02-10SHANGHAI LAMSHINE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610044215.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-02-10
Estimated Expiration
2046-01-14

AI Technical Summary

Technical Problem

In indoor environments, the visual positioning of micro drones is easily affected by repetitive textured scenes, leading to the accumulation of positioning errors and even collision accidents.

Method used

An ultrasonic environmental perception method is adopted, which combines a visual camera, an inertial measurement unit (IMU), and a time-of-flight (ToF) ranging sensor. Repeated texture scenes are identified through frequency domain transformation, and inter-frame displacement and corridor width change sequences of image sequences are extracted. Multi-source data fusion positioning is performed, and the positioning results are optimized using an extended Kalman filter model.

Benefits of technology

It achieves intelligent recognition and localization correction for repetitive texture scenes, improves the accuracy and stability of localization, solves the drift problem of visual localization in repetitive texture environments, and ensures the safe operation of drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121498718A_ABST
    Figure CN121498718A_ABST
Patent Text Reader

Abstract

The invention discloses an ultrasonic environment perception miniature unmanned aerial vehicle indoor positioning method and system, and relates to the field of non-electrical variable control or adjustment systems, and the method comprises the steps: the system collects an image sequence, IMU and ToF data; and judging a repeated texture scene by analyzing image frequency domain characteristics and wall distance changes. In the scene, displacement between adjacent image frames is calculated and accumulated to obtain visual displacement. And matching a corridor width change sequence measured by ToF with a map reference sequence to obtain a longitudinal position and a transverse offset. And finally, fusing the visual displacement, the position information and the IMU data through a filter to obtain a final positioning result. By implementing the method, the accuracy of unmanned aerial vehicle positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of control or regulation systems for non-electrical variables, and more particularly to an indoor positioning method and system for a micro unmanned aerial vehicle (UAV) with ultrasonic environmental perception. Background Technology

[0002] With the rapid development of drone technology, the application of micro drones in indoor scenarios is becoming increasingly widespread. Accurate indoor positioning capability is a crucial foundation for micro drones to complete tasks such as indoor inspection and material delivery. Since GPS signals are unavailable indoors, other positioning methods are needed to achieve autonomous navigation and positioning for micro drones.

[0003] Currently, indoor drone positioning primarily employs visual positioning schemes. This involves acquiring image sequences using a visual camera, extracting image feature points, and tracking the movement of these feature points to estimate the drone's position changes. To further improve positioning accuracy, an inertial measurement unit (IMU) is introduced to acquire the drone's attitude and motion data, which is then fused with the visual positioning results.

[0004] However, in indoor corridors and other scenarios, walls often have repetitive textures and patterns, such as regularly arranged bricks or decorative stripes. Such repetitive textures can lead to a large number of mismatches in visual feature point extraction and tracking, resulting in cumulative errors in vision-based position estimation. Ultimately, this can cause drones to deviate from their intended flight paths or even collide with other drones. Summary of the Invention

[0005] This application provides an indoor positioning method and system for micro unmanned aerial vehicles (UAVs) based on ultrasonic environmental perception, which can improve the accuracy of UAV positioning.

[0006] Firstly, this application provides an indoor positioning method for micro unmanned aerial vehicles (UAVs) based on ultrasonic environmental perception, applied to an indoor positioning system for micro UAVs. The method includes: acquiring image sequence data from a visual camera, inertial motion data from an inertial measurement unit (IMU), and distance measurement data from a time-of-flight (ToF) sensor; performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks; measuring the distances to the left and right side walls using the ToF sensor; identifying the current scene as a repeating texture scene when the variation amplitude of the periodic spectral peaks and the distances to the left and right side walls is within a preset threshold range; and extracting the time-of-flight distances between adjacent image frames in the image sequence data within the repeating texture scene. The inter-frame displacement transformation is accumulated along the flight direction to obtain the visual cumulative displacement. The distances to the left and right side walls are continuously measured at multiple measurement positions along the flight direction using a Time-of-Flight (ToF) sensor. These side wall distances are then sequentially arranged to form a corridor width variation sequence. This corridor width variation sequence is matched with a reference width variation sequence in a pre-stored map. The longitudinal positioning position is determined based on the location of the matched reference width variation sequence on the map, and the lateral offset is calculated based on the distances to the left and right side walls. The visual cumulative displacement, longitudinal positioning position, lateral offset, and inertial motion data are input to a filter and fused to output the fused positioning result.

[0007] In the above embodiments, the system achieves intelligent recognition and localization correction capabilities for repetitive texture scenes. By integrating multi-source data from visual cameras, IMUs, and ToF sensors, the system can perform real-time recognition and feature extraction of repetitive texture scenes, calculate and correct the cumulative deviation of visual localization. Simultaneously, based on the corridor width variation sequence measured by the ToF sensor, position matching is performed, and combined with motion information provided by the IMU, a reliable localization solution model is established, solving the problem of localization drift in repetitive texture environments using traditional visual localization.

[0008] In conjunction with some embodiments of the first aspect, in some embodiments, the step of performing frequency domain transformation on image sequence data to obtain periodic spectral peaks, measuring the distances to the left and right side walls using a time-of-flight ranging sensor (ToF), and identifying the current scene as a repetitive texture scene when the variation amplitude of the periodic spectral peaks and the distances to the left and right side walls is within a preset threshold range specifically includes: selecting a preset number of image frames from the image sequence data to form an image analysis window; extracting feature points from each image frame in the image analysis window; calculating the number of feature point matches between adjacent image frames; when the number of feature point matches is greater than a first preset threshold, performing a two-dimensional fast Fourier transform on the image frames in the image analysis window; extracting the horizontal and vertical spectral components in the transformed frequency domain space; calculating the peak value of the spectral components as the periodic spectral peak; obtaining the left and right side wall distance measurement sequence within a preset time period from the time-of-flight ranging sensor (ToF); calculating the root mean square error of the left and right side wall distance measurement sequence as the variation amplitude of the left and right side wall distances; and determining the current scene as a repetitive texture scene when the periodic spectral peak is greater than a second preset threshold and the variation amplitude of the left and right side wall distances is less than a third preset threshold.

[0009] In the above embodiments, the system constructs a fast and accurate identification mechanism for repetitive texture scenes. Combining frequency domain analysis of image sequences and ToF distance measurement data, the system establishes a scene recognition model based on multi-feature fusion. By extracting the periodic spectral features of the image and the wall distance variation features, the system can accurately determine whether there are repetitive textures in the current scene. Simultaneously, an adaptive threshold judgment method is adopted to improve the robustness of scene recognition and provide a reliable basis for subsequent localization strategy adjustments.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, the step of performing sequence feature matching between the corridor width change sequence and a reference width change sequence in a pre-stored map, and determining the longitudinal positioning position based on the position of the matched reference width change sequence in the map, specifically includes: comparing the corridor width change sequence and the reference width change sequence through a sliding window, calculating the sequence correlation coefficient for each sliding window position; obtaining the sliding window position with the largest correlation coefficient, and using the starting position of the reference width change sequence corresponding to the sliding window position as the longitudinal positioning position; calculating the distance from the current position to the corridor centerline based on the distance between the left and right side walls, and using the distance as the lateral offset.

[0011] In the above embodiments, the system achieves precise positioning based on corridor width variations. By continuously ranging the left and right side walls of the corridor using a ToF sensor, the system acquires a sequence of corridor width variations reflecting environmental characteristics. Through sliding window matching with a reference sequence in a pre-stored map, the sequence correlation is calculated to determine the precise position of the UAV. Combined with lateral offset calculation, the system establishes a complete position estimation model, providing accurate position feedback for navigation control.

[0012] In conjunction with some embodiments of the first aspect, in some embodiments, the step of fusing visual cumulative displacement, longitudinal positioning position, lateral offset, and inertial motion data into a filter and outputting a fused positioning result specifically includes: extracting acceleration and angular velocity components from the inertial motion data; performing zero-bias correction and integration on the acceleration and angular velocity components to obtain inertial position prediction values; constructing a state vector, including position, velocity, and attitude components; constructing a position observation equation using the inertial position prediction value, visual cumulative displacement, and longitudinal positioning position; constructing a lateral constraint equation using the lateral offset; constructing an extended Kalman filter model based on the state vector, position observation equation, and lateral constraint equation; and iteratively updating the state vector using the extended Kalman filter model to obtain the fused positioning result.

[0013] In the above embodiments, the system achieves deep fusion of multi-sensor data. Based on the extended Kalman filter framework, the system constructs a state estimation model that includes position, velocity, and attitude. By optimizing and fusing visual cumulative displacement, IMU integral position, and ToF geometric localization results, a compensation mechanism for measurement errors of each sensor is established, improving the accuracy and reliability of the localization results.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, after fusing visual cumulative displacement, longitudinal positioning position, lateral offset, and inertial motion data into a filter and outputting a fused positioning result, the method further includes: acquiring distance measurement data of a Time-of-Flight (ToF) sensor in the horizontal plane, determining the width and spacing of shelf aisles in a repetitive texture scene; extracting shelf edge features from image sequence data, and determining the extension direction of the shelf aisles by combining the distance measurement data from the ToF sensor; when the ToF sensor detects an obstacle ahead, calculating the position of adjacent aisles based on the width and spacing of the shelf aisles; controlling the UAV to turn to adjacent aisles based on visual cumulative displacement, maintaining the relative position with the shelf using shelf edge features during the turning process; and recording the position information of the shelf aisle where the obstacle is located into a pre-stored map.

[0015] In the above embodiments, the system has constructed a complete intelligent obstacle avoidance function. Based on ToF horizontal plane scanning and visual feature recognition, the system achieves three-dimensional perception of environmental obstacles. By analyzing the geometric features and spatial distribution of the shelf aisle, the system can calculate alternative obstacle avoidance paths and, combined with visual feature tracking, achieve smooth turning and obstacle avoidance, ensuring the safe operation of the UAV in complex environments.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, after recording the location information of the shelf aisle where the obstacle is located into a pre-stored map, the method further includes: calculating the width difference between two adjacent shelf aisles based on distance data measured by a time-of-flight ranging sensor (ToF); when the width difference is greater than a preset width threshold, determining that the shelf with the smaller width among the two adjacent shelf aisles is temporarily occupied; extracting the spatial distribution features of the obstacle from image sequence data and distance measurement data from the time-of-flight ranging sensor (ToF); when a change in the spatial distribution features is detected, marking the changed area as a dynamic obstacle avoidance area; and during the execution of the obstacle avoidance trajectory, selecting the shelf aisle that is not marked as a temporarily occupied aisle or a dynamic obstacle avoidance area as an alternative obstacle avoidance path.

[0017] In the above embodiments, the system establishes a dynamic environment perception and processing mechanism. By analyzing the width differences between adjacent shelf aisles and the spatiotemporal distribution characteristics of obstacles, the system can identify temporarily occupied aisles and dynamically changing areas. Based on multi-level environmental feature analysis, the system realizes real-time updates of obstacle avoidance strategies, enhancing the adaptability and safety of UAVs in dynamic environments.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, after fusing visual cumulative displacement, longitudinal positioning position, lateral offset, and inertial motion data into a filter and outputting a fused positioning result, the method further includes: extracting the angular velocity component from the inertial motion data; identifying that the UAV is about to enter a corner scene when the rate of change of the angular velocity component exceeds a preset angular velocity threshold; within a preset time window before entering the corner scene, adjusting the fusion weight of visual positioning from a first weight value to a second weight value through a linear gradient function, while adjusting the fusion weight of inertial positioning from a third weight value to a fourth weight value, and adjusting the fusion weight of ToF-based geometric positioning from a fifth weight value to a sixth weight value, wherein the second weight value is less than the first weight value, the fourth weight value is greater than the third weight value, and the sixth weight value is greater than the fifth weight value; recording the scene type identifier and corridor width change sequence at the corresponding position of the grid cell in the pre-stored map grid cell, wherein the scene type identifier includes a repeating texture scene identifier and a corner scene identifier; during flight, querying the grid cells within a preset distance range ahead based on the current position and flight direction to obtain the scene type identifier ahead, and performing sensor fusion weight adjustment for the corresponding scene in advance based on the scene type identifier ahead.

[0019] In the above embodiments, the system implements a scene-aware sensor fusion mechanism. By analyzing IMU angular velocity data, the system can identify corner scenes in advance and adjust the fusion weights of different sensors through a linear gradient function. Combined with scene type information in the pre-stored map, the system achieves predictive adjustment of the sensor fusion strategy and establishes a complete scene-adaptive localization framework.

[0020] In a second aspect, embodiments of this application provide a micro unmanned aerial vehicle (UAV) indoor positioning system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, which includes computer instructions, and the one or more processors call the computer instructions to cause the micro unmanned aerial vehicle (UAV) indoor positioning system to perform the methods described in the first aspect and any possible implementation thereof.

[0021] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a micro unmanned aerial vehicle (UAV) indoor positioning system, cause the micro UAV indoor positioning system to perform the method described in the first aspect and any possible implementation thereof.

[0022] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a micro unmanned aerial vehicle (UAV) indoor positioning system, cause the micro UAV indoor positioning system to perform the method described in the first aspect and any possible implementation thereof.

[0023] Understandably, the micro-UAV indoor positioning system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.

[0024] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By employing a multi-sensor fusion indoor positioning method, the complementary advantages of visual, IMU, and ToF sensors can be fully utilized to achieve intelligent recognition and positioning correction in repetitive texture scenes. Position matching was performed using corridor width variation sequences measured by the ToF sensor, establishing a positioning correction mechanism based on environmental geometric features. Simultaneously, by combining motion information provided by the IMU and cumulative visual displacement, a multi-source data fusion model was constructed, effectively solving the problem of cumulative errors easily generated in repetitive texture environments by single-vision positioning in existing technologies. This resulted in stable and reliable indoor positioning functionality, providing accurate position information support for the autonomous navigation of UAVs.

[0025] 2. By employing a scene recognition method based on frequency domain analysis and Time-of-Flight (ToF) distance measurement, periodic spectral features can be extracted from image sequences. Combined with ToF-measured wall distance variations, a multi-dimensional scene feature analysis model is constructed. A complete process for recognizing repetitive texture scenes is established through setting multiple preset thresholds for feature matching and judgment. The system can assess current environmental features in real time and adjust its positioning strategy accordingly. This effectively solves the problem of inappropriate positioning strategy selection due to the difficulty in recognizing complex environmental features in existing technologies. Thus, it achieves adaptive adjustment of environmental perception and positioning strategy, ensuring the positioning performance of the UAV in different scenarios.

[0026] 3. By employing a location matching method based on corridor width variation sequences, the geometric features of the environment can be fully utilized for positioning correction. Continuous ranging at multiple locations using a ToF sensor acquires width variation sequences reflecting local environmental characteristics. A sliding window sequence matching algorithm is used to calculate the correlation with reference sequences in a pre-stored map, achieving large-scale positioning based on environmental features. By analyzing the distances to the left and right side walls to calculate lateral offset, the problem of positioning drift easily occurring in textured, repetitive areas in existing technologies is effectively solved, thus achieving a complete position estimation that takes into account both longitudinal positioning and lateral constraints. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating an indoor positioning method for a micro unmanned aerial vehicle using ultrasonic environmental perception, as described in this application. Figure 2 This is another flowchart illustrating the ultrasonic environmental perception-based indoor positioning method for micro-UAVs in this application embodiment; Figure 3 This is a schematic diagram of the physical device structure of a micro unmanned aerial vehicle indoor positioning system in the embodiments of this application. Detailed Implementation

[0028] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.

[0029] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0030] To facilitate understanding, the application scenarios of the embodiments of this application are described below.

[0031] To facilitate understanding, the method provided in this implementation will be described in detail below, using the above scenario as an example. Please refer to [link / reference]. Figure 1 This is a flowchart illustrating an indoor positioning method for a micro unmanned aerial vehicle (UAV) based on ultrasonic environmental perception, as described in this application.

[0032] S101: Acquire image sequence data from the visual camera, inertial motion data from the inertial measurement unit (IMU), and distance measurement data from the time-of-flight (ToF) sensor.

[0033] Among them, image sequence data represents a data stream composed of multiple frames of images continuously acquired by a vision camera at a fixed frequency; inertial measurement unit (IMU) refers to a sensor device used to measure the acceleration and angular velocity of an object; and time-of-flight (ToF) range sensor refers to a sensor device that measures distance by emitting and receiving light signals.

[0034] This step is performed when the drone takes off and begins its indoor positioning task. Specifically, the system acquires RGB image sequences at a frequency of no less than 30 frames per second using a vision camera, simultaneously acquires triaxial acceleration and triaxial angular velocity data output by the IMU at a frequency of over 200Hz, and acquires distance measurement data output by the ToF sensor at a frequency of no less than 10Hz, thus achieving synchronous acquisition of multi-sensor data.

[0035] In some embodiments, the acquisition and synchronization of multi-sensor data can be achieved in several ways: Optionally, hardware trigger signals can be used to synchronize the sampling clocks of the three sensors, and the sampled data can be buffered into three FIFO queues respectively, with data alignment and fusion performed based on timestamp information; alternatively, software timers can be used to trigger sensor sampling, and time interpolation and data synchronization can be performed in the data processing thread. It is understood that other sensor synchronization schemes can also be used to achieve data acquisition.

[0036] S102. Perform frequency domain transformation on the image sequence data to obtain periodic spectral peaks. Use a time-of-flight ranging sensor (ToF) to measure the distance between the left and right side walls. When the variation amplitude of the periodic spectral peaks and the distance between the left and right side walls is within a preset threshold range, identify the current scene as a repeating texture scene.

[0037] Among them, frequency domain transformation refers to the mathematical transformation method of converting time domain signals to frequency domain; periodic spectral peaks refer to the amplitude of the periodic characteristics exhibited by the signal in the frequency domain; variation amplitude represents the intensity of signal fluctuations; and repetitive texture scenes refer to scenes in the environment where there are regular repetitive patterns.

[0038] This step is performed after data acquisition is complete. Specifically, the system first performs a Fast Fourier Transform on the acquired image sequence to analyze the spectral characteristics of the images in the horizontal and vertical directions and extract the peak values ​​of the main frequency components. Simultaneously, a ToF sensor is used to measure the distances to the left and right walls, and the statistical characteristics of the distance values ​​are calculated. When both the spectral peak values ​​and the distance variation characteristics meet preset conditions, it is determined that the current scene is a repeating texture scene.

[0039] In some embodiments, scene recognition can be achieved in several ways: Optionally, a sliding window approach can be used to process the image sequence, performing a two-dimensional FFT transform on each window to extract the spectral peaks in the horizontal and vertical directions, and combining the variance features of the ToF distance data to determine the scene; alternatively, a deep learning method can be used to train a scene classification model, taking the spectral features and distance features as input and outputting the scene type probability. It is understood that other methods can also be used to achieve scene recognition.

[0040] S103. In a scene with repetitive textures, extract the inter-frame displacement transformation between adjacent image frames in the image sequence data, and accumulate the inter-frame displacement transformation along the flight direction to obtain the visual cumulative displacement.

[0041] Among them, the inter-frame displacement transformation represents the relative motion displacement between adjacent image frames; the visual cumulative displacement refers to the total displacement accumulated along the direction of motion; and the flight direction refers to the forward direction of the UAV.

[0042] This step is performed after identifying repetitive texture scenes. Specifically, the system extracts and matches feature points from adjacent frames in the image sequence, calculates the inter-frame homography matrix, and decomposes it to obtain relative motion parameters. The inter-frame displacements are accumulated along the flight direction to obtain a vision-based cumulative displacement estimate.

[0043] In some embodiments, visual displacement estimation can be achieved in several ways: optionally, feature points can be extracted using the FAST corner detection algorithm, the motion of the feature points can be tracked using optical flow, and inter-frame motion parameters can be calculated; optionally, the fundamental matrix can be estimated using a feature point matching method, and the camera motion can be decomposed to obtain the motion. It is understood that other visual odometry methods can also be used to achieve displacement estimation.

[0044] S104. Based on the time-of-flight ranging sensor (ToF), continuously measure the distances between the left and right side walls at multiple measurement locations along the flight direction, and arrange the side wall distances at multiple measurement locations into a sequence of corridor width changes.

[0045] Among them, the measurement position refers to the spatial sampling point where the ToF sensor performs distance measurement; the distance between the left and right side walls refers to the actual distance value from the sensor to the left and right side walls of the corridor; the corridor width change sequence represents the ordered data set of corridor width at multiple sampling locations along the flight direction; continuous measurement refers to the distance measurement process that is coherent in time and space.

[0046] This step is executed immediately after the sensor data is acquired. Specifically, the system controls the ToF sensor to sample along the flight path at fixed spatial intervals (e.g., 0.5 meters), measuring the distances to the left and right walls at each sampling location. The measured left and right distance values ​​are summed to obtain the corridor width at that location, and the width values ​​from N consecutive locations (e.g., N=20) are arranged into a sequence in chronological order. This sequence reflects the local variation characteristics of the corridor's geometry and can be used for subsequent location matching.

[0047] In some embodiments, the corridor width sequence can be constructed in several ways: optionally, multiple ToF sensors can be used to simultaneously measure distances in different directions, and accurate wall distances can be obtained through coordinate transformation and data fusion; optionally, attitude compensation can be performed on the ToF measurements using IMU data to eliminate measurement errors caused by changes in flight attitude; optionally, Kalman filtering can be used to smooth the raw measurement data to improve the reliability of the sequence data. It is understood that other methods can also be used to obtain the corridor width sequence.

[0048] It should be noted that the "corridor width variation sequence" described in this application refers not only to the gradual change in the macroscopic physical width of the corridor, but also includes detailed information on the local geometric features of the walls on both sides of the corridor. In actual indoor environments, even corridors that are macroscopically of uniform width (such as office building corridors with a constant width of 2 meters) are not absolutely flat planes, but rather have structures with significant geometric features.

[0049] Specifically, the features relied upon for feature matching include, but are not limited to: Incremental features (recesses): such as recesses in doorways on both sides of a corridor, open doors, fire hydrant installation slots, elevator entrances, etc. These features are manifested in ToF measurement data as an instantaneous step increase in width (pulse peak). Reduction features (protrusions): such as load-bearing columns, wall decoration protrusions, exposed pipe wells, etc. These features are manifested as an instantaneous reduction in width (pulse trough) in ToF measurement data. Texture features: abrupt changes in reflectivity caused by variations in wall material (requires ToF sensor support for intensity echo).

[0050] Therefore, the "corridor width variation sequence" constructed in this embodiment is actually a set of "one-dimensional depth fingerprints" reflecting the geometric structure of the environment. Even in a long, straight corridor of equal width, as long as any of the above-mentioned geometric features exist, the sequence is unique, thus enabling the determination of the longitudinal position through feature matching.

[0051] S105. Perform sequence feature matching between the corridor width change sequence and the reference width change sequence in the pre-stored map. Determine the longitudinal positioning position based on the position of the matched reference width change sequence in the map. Calculate the lateral offset based on the distance between the left and right side walls.

[0052] Among them, the pre-stored map refers to a pre-established database containing environmental geometric information. The pre-stored map not only records the width of the corridor, but also marks the position coordinates and size information of key geometric features (such as doors and pillars); the reference width change sequence refers to the standard corridor width change pattern recorded in the map; sequence feature matching refers to calculating the degree of matching between two sequences; the vertical positioning position represents the position coordinates along the corridor direction; and the lateral offset represents the vertical distance to the center line of the corridor.

[0053] This step is performed after obtaining the corridor width change sequence. Specifically, the system performs a sliding window matching between the currently obtained width change sequence and all reference sequences stored in the map database, calculating the normalized cross-correlation coefficient. The matching result with the highest correlation coefficient exceeding the threshold is selected, and the starting position of the corresponding reference sequence on the map is taken as the current vertical position. At the same time, based on the difference in distance between the left and right walls, the lateral offset distance of the drone relative to the corridor centerline is calculated.

[0054] In some embodiments, location matching and offset calculation can be achieved in several ways: optionally, a dynamic time warping (DTW) algorithm can be used for sequence matching to adapt to the nonlinear characteristics of corridor width variations; optionally, multi-scale sequence templates can be constructed to improve positioning efficiency and robustness through hierarchical matching; optionally, wall material and texture features can be combined to assist sequence matching and improve the accuracy of location identification. It is understood that other methods can also be used to achieve location matching.

[0055] S106. Input the visual cumulative displacement, longitudinal positioning position, lateral offset and inertial motion data into the filter and fuse them to output the fused positioning result.

[0056] Here, the filter represents the state estimator used for multi-source data fusion; the fusion positioning result refers to the optimal position estimate obtained by integrating information from various sensors; and the inertial motion data includes acceleration and angular velocity measurements.

[0057] This step is performed after obtaining various position estimates. Specifically, the system constructs an extended Kalman filter, using visual displacement, Time-of-Flight (ToF) based position matching results, and IMU integral position as observations, and establishes state transition equations and observation equations. Through iterative prediction and update processes, the advantages of data from various sensors are integrated, their respective cumulative errors are suppressed, and finally, a reliable position estimate is output.

[0058] In some embodiments, multi-sensor data fusion can be achieved in several ways: optionally, an adaptive weighting method based on the covariance matrix can be used to dynamically adjust the fusion weights according to the reliability of each sensor's data; optionally, particle filtering can be introduced to process the nonlinear observation model and improve positioning accuracy; optionally, filter parameters can be adjusted in conjunction with scene recognition results to adapt to positioning requirements in different environments. It is understood that other methods can also be used to achieve data fusion.

[0059] The following provides a more detailed description of the process of the method provided in this implementation. Please refer to [link / reference]. Figure 2 This is another flowchart illustrating the ultrasonic environmental perception-based indoor positioning method for micro-UAVs in this application.

[0060] S201. Acquire image sequence data from the visual camera, inertial motion data from the inertial measurement unit (IMU), and distance measurement data from the time-of-flight (ToF) sensor.

[0061] In some embodiments, this step is similar to step S101, and will not be described again here.

[0062] S202. Select a preset number of image frames from the image sequence data to form an image analysis window. Extract feature points from each image frame in the image analysis window and calculate the number of feature point matches between adjacent image frames.

[0063] Among them, the preset number represents a fixed number of image frames, usually 8-12 frames are selected to ensure the continuity of analysis time and computational efficiency; the image analysis window refers to the set of continuous image frames used for feature extraction and motion analysis, reflecting the changing characteristics of the scene in a short period of time; feature points represent salient corner points or edge points in the image, which have the characteristic of obvious gray-scale changes in local areas; the feature point matching number refers to the number of feature point pairs successfully matched between two image frames, reflecting the similarity of image content and motion continuity.

[0064] The system selects 10 consecutive images from the image sequence cache in real time as the analysis window, and uses the FAST corner detection algorithm to extract feature points in each frame. The algorithm first selects candidate pixels in the image and compares the grayscale value difference between the point and its 16 surrounding pixels. When the difference between the grayscale values ​​of N consecutive surrounding pixels (N>9) and the center point exceeds a threshold, the point is marked as a corner. For each detected corner, its BRIEF binary descriptor is calculated, obtained by comparing the grayscale values ​​of several randomly selected pixel pairs in the corner's neighborhood. Hamming distance is calculated for the feature point descriptors of adjacent frames, and the nearest neighbor ratio rule (threshold set to 0.7) is used to filter matching pairs and obtain initial matching results. Further, the RANSAC algorithm is used to eliminate outliers, obtaining the final valid matching point pairs. For 640×480 resolution images, 300-500 feature points are typically extracted per frame, and 100-200 stable matching pairs can be obtained between adjacent frames. The system records the number of these matching pairs for subsequent scene feature analysis.

[0065] S203. When the number of matching feature points is greater than the first preset threshold, perform a two-dimensional fast Fourier transform on the image frame in the image analysis window, extract the horizontal and vertical spectral components in the transformed frequency domain space, and calculate the peak value of the spectral components as the periodic spectral peak value.

[0066] Among them, the two-dimensional fast Fourier transform is an efficient mathematical operation method for converting two-dimensional image signals from the spatial domain to the frequency domain, which can reveal the periodic structural features of the image; the frequency domain space represents a two-dimensional coordinate system of signal frequency distribution, with the horizontal and vertical axes corresponding to the spatial frequencies in the horizontal and vertical directions, respectively; the spectral components refer to the energy distribution of the signal at different frequencies, and their amplitudes reflect the intensity of the corresponding frequency components; the periodic spectral peaks represent significant local maxima in the spectrum, corresponding to the main periodic features in the image.

[0067] When the number of matching feature points exceeds a set threshold (e.g., 150 pairs), the system performs the following processing on each frame of the image in the analysis window: First, the image is converted to grayscale and histogram equalization is performed to enhance image contrast. Then, a two-dimensional FFT transformation is performed on the preprocessed image to obtain a complex matrix in the frequency domain. The modulus of the complex matrix is ​​calculated to obtain the power spectrum, and a logarithmic transformation is performed to enhance the visual effect. The power spectrum is accumulated along the horizontal (u-axis) and vertical (v-axis) directions to obtain a one-dimensional spectrum curve. A sliding window method (window size is 1 / 10 of the total spectrum length) is used to detect peaks in the spectrum curve, and local maxima points within the window that are greater than twice the mean are recorded. The detected peaks are sorted by amplitude, and the largest peak is selected as the quantification index of the periodic feature. For example, for an image with a regular brick wall texture, significant spectral peaks will appear at the spatial frequencies corresponding to the brick spacing, and the peak amplitude is usually 3-5 times the mean of the background spectrum.

[0068] S204. Obtain the left and right side wall distance measurement sequence within a preset time period from the Time-of-Flight (ToF) ranging sensor, and calculate the mean square error of the left and right side wall distance measurement sequence as the change range of the left and right side wall distance.

[0069] Among them, the two-dimensional fast Fourier transform is an efficient mathematical operation method for converting two-dimensional image signals from the spatial domain to the frequency domain, which can reveal the periodic structural features of the image; the frequency domain space represents a two-dimensional coordinate system of signal frequency distribution, with the horizontal and vertical axes corresponding to the spatial frequencies in the horizontal and vertical directions, respectively; the spectral components refer to the energy distribution of the signal at different frequencies, and their amplitudes reflect the intensity of the corresponding frequency components; the periodic spectral peaks represent significant local maxima in the spectrum, corresponding to the main periodic features in the image.

[0070] When the number of matching feature points exceeds a set threshold (e.g., 150 pairs), the system performs the following processing on each frame of the image in the analysis window: First, the image is converted to grayscale and histogram equalization is performed to enhance image contrast. Then, a two-dimensional FFT transformation is performed on the preprocessed image to obtain a complex matrix in the frequency domain. The modulus of the complex matrix is ​​calculated to obtain the power spectrum, and a logarithmic transformation is performed to enhance the visual effect. The power spectrum is accumulated along the horizontal (u-axis) and vertical (v-axis) directions to obtain a one-dimensional spectrum curve. A sliding window method (window size is 1 / 10 of the total spectrum length) is used to detect peaks in the spectrum curve, and local maxima points within the window that are greater than twice the mean are recorded. The detected peaks are sorted by amplitude, and the largest peak is selected as the quantification index of the periodic feature. For example, for an image with a regular brick wall texture, significant spectral peaks will appear at the spatial frequencies corresponding to the brick spacing, and the peak amplitude is usually 3-5 times the mean of the background spectrum.

[0071] S205. When the peak value of the periodic spectrum is greater than the second preset threshold and the change in the distance between the left and right side walls is less than the third preset threshold, the current scene is determined to be a repeating texture scene.

[0072] Among them, the preset duration represents a fixed data acquisition time window, usually selected as 1-3 seconds to ensure the reliability of data statistics; the distance measurement sequence refers to the set of ToF sensor distance data acquired continuously within the sampling time; the root mean square error represents a statistical indicator of the degree of data fluctuation, reflecting the stability of the wall distance.

[0073] Within a preset 2-second time window, the system controls the ToF sensor to simultaneously measure the distances to the left and right walls at a sampling frequency of 50Hz. For each measurement, the sensor emits a modulated near-infrared light pulse, and the distance value is calculated by detecting the phase delay of the reflected light signal. Approximately 100 sampling points {L1, L2, ..., L...} are obtained for the left wall distance within 2 seconds. 100} and the distance from the sampling points {R1, R2, ..., R} to the right wall 100 The arithmetic mean μL and μR of the two sequences are calculated respectively. Then, the sum of squares of the differences between each sampling point and the mean is calculated, divided by the number of sampling points, and the square root is taken to obtain the root mean square deviations σL and σR of the left and right wall distance sequences. The specific calculation formulas are: σL = sqrt(Σ(Li-μL)² / 100) and σR = sqrt(Σ(Ri-μR)² / 100). The larger root mean square deviation value max(σL, σR) is selected as a measure of the variation in wall distance. When this value is less than a preset threshold (e.g., 5cm), it indicates that the wall distance is relatively stable, which helps to determine that the current environment is a regular corridor. At the same time, the system also records the absolute values ​​of the left and right wall distances for subsequent location estimation and path planning.

[0074] S206. In a scene with repetitive textures, extract the inter-frame displacement transformation between adjacent image frames in the image sequence data, and accumulate the inter-frame displacement transformation along the flight direction to obtain the visual cumulative displacement.

[0075] In some embodiments, this step is similar to step S103, and will not be described again here.

[0076] S207. Based on the time-of-flight ranging sensor (ToF), the distances between the left and right side walls are continuously measured at multiple measurement positions along the flight direction, and the side wall distances at multiple measurement positions are sequentially arranged into a corridor width variation sequence.

[0077] In some embodiments, this step is similar to step S104, and will not be described again here.

[0078] S208. Compare the corridor width change sequence with the reference width change sequence using a sliding window, and calculate the sequence correlation coefficient at each sliding window position.

[0079] S209. Obtain the position of the sliding window with the largest correlation coefficient, and take the starting position of the reference width change sequence corresponding to the sliding window position as the vertical positioning position.

[0080] S210. Calculate the distance from the current position to the central axis of the corridor based on the distance between the left and right side walls, and use the distance as the lateral offset. Calculate the lateral offset based on the distance between the left and right side walls.

[0081] Among them, the sliding window represents a fixed-length data segment that moves along the sequence and is used for local sequence alignment; the sequence correlation coefficient is a statistic that measures the similarity between two data sequences, with a value range of [-1, 1]; the longitudinal positioning position refers to the absolute position coordinates of the UAV in the direction of the corridor; the corridor centerline refers to the center line of the walls on both sides of the corridor; and the lateral offset represents the vertical distance of the UAV from the corridor centerline.

[0082] The system first sets the sliding window length W (e.g., 20 sampling points), and then processes the currently obtained corridor width variation sequence {w1, w2, ..., w...}. n} and the reference sequence {r1, r2, ..., r} in the pre-stored map m A sliding comparison is performed. For each window position i, the Pearson correlation coefficient of the sequence segments within the window is calculated: ρᵢ = cov(W_i, R_i) / [std(W_i)×std(R_i)], where W_i and R_i are the data segments of the two sequences at the current window position, cov represents the covariance, and std represents the standard deviation. All possible window positions are iterated to find the position i_max with the highest correlation coefficient. The starting coordinates of the reference sequence corresponding to the i_max position are taken as the longitudinal position of the UAV. Simultaneously, based on the currently measured distances L and R between the left and right walls, the lateral offset is calculated: offset = (LR) / 2. This value represents the distance the UAV deviates from the corridor centerline; a positive value indicates a deviation to the right, and a negative value indicates a deviation to the left. This yields the complete two-dimensional position of the UAV in the corridor coordinate system.

[0083] Furthermore, for the special scenario of "equal width corridors", the process of matching the corridor width change sequence with the reference width change sequence in the pre-stored map preferably adopts a local matching strategy based on salient features.

[0084] The specific steps are as follows: The acquired real-time width sequence is preprocessed, and the first derivative or gradient of the sequence is calculated to identify abrupt change points (i.e., geometric feature points) in the data. For example, when a drone flies over a doorway, the ToF measurement of the distance to the left or right will generate a rectangular wave signal exhibiting a "distance increase-hold-decrease" pattern. The system extracts the feature attributes of these abrupt change points, including: feature type (recess or protrusion), feature width (duration along the flight direction), feature depth (depth of the recess or height of the protrusion), and the spacing between adjacent features. The extracted real-time feature attribute set is then topologically matched with the reference feature attribute set recorded in a pre-stored map. For example, if the real-time sequence detects "a recess (doorway) with a width of 0.9 meters and a depth of 0.2 meters, and a protrusion (pillar) with a width of 0.5 meters detected 3 meters in front of it," the system will search the map database for areas with the same topological structure. In this way, even in extreme environments where visual textures are highly repetitive (such as white walls) and the overall width of the corridor is constant, the system can still accurately lock the longitudinal position of the drone by utilizing the distribution patterns of geometric features such as doorways and pillars, effectively solving the drift problem of traditional visual odometry in long corridor scenarios.

[0085] S211. Extract acceleration and angular velocity components from inertial motion data, perform zero-bias correction and integration on acceleration and angular velocity components, and obtain the predicted inertial position value.

[0086] Among them, the acceleration component refers to the three-axis linear acceleration measured by the IMU; the angular velocity component is the three-axis angular rate measured by the IMU; zero bias correction represents the process of eliminating the static error of the sensor; integration operation refers to the time integration of acceleration and angular velocity data.

[0087] The system extracts triaxial acceleration [ax, ay, az] and angular velocity [ωx, ωy, ωz] from the raw IMU data. First, zero-bias correction is performed: 5 seconds of data are collected while the system is stationary, and the average values ​​of each axis are calculated as zero biases [bax, bay, baz] and [bωx, bωy, bωz]. These zero-bias values ​​are subtracted from the real-time data. The corrected angular velocity is integrated in first order to obtain the attitude angle change, and the rotation matrix R from the computational volume coordinate system to the world coordinate system is updated using quaternions. The acceleration data is transformed to the world coordinate system using R, and after subtracting the gravitational acceleration g, a double integration is performed: the first integration yields the velocity v(t) = ∫(a(t) - g)dt, and the second integration yields the position p(t) = ∫v(t)dt. The integration is implemented using the trapezoidal rule, with the integration time step set to the IMU sampling period (e.g., 0.005 seconds). The resulting p(t) is the predicted position value based on the inertial data.

[0088] S212. Construct a state vector, including position, velocity, and attitude components, and construct a position observation equation using the inertial position prediction, visual cumulative displacement, and longitudinal positioning position.

[0089] Here, the state vector is a set of variables describing the dynamic characteristics of the system; the position component contains three-dimensional spatial coordinates [x, y, z]; the velocity component represents the motion speed in three directions [vx, vy, vz]; the attitude component is represented by Euler angles [roll, pitch, yaw]; and the position observation equation describes the functional relationship between the measured values ​​and the state variables.

[0090] The system constructs a 13-dimensional state vector X = [x, y, z, vx, vy, vz, roll, pitch, yaw, bax, bay, baz, bω], which includes position, velocity, attitude angle, and IMU zero-bias term. Based on the state vector, the nonlinear system state equation is constructed: X(k+1) = f(X(k), u(k)) + w(k), where u(k) is the IMU input and w(k) is the process noise. The position observation equation adopts a multi-source fusion form: Z(k) = [Zimu(k); Zvis(k); Ztof(k)], corresponding to the inertial position prediction, visual cumulative displacement, and ToF-based longitudinal position measurement, respectively. The measurement noise covariance matrix R of the observation equation is set according to the measurement accuracy of each sensor.

[0091] S213. Construct lateral constraint equations using lateral offsets, and build an extended Kalman filter model based on the state vector, position observation equations, and lateral constraint equations.

[0092] Among them, the lateral constraint equation represents the geometric constraints of the UAV's lateral movement in the corridor; the extended Kalman filter model is a state estimator for handling nonlinear systems; and the position observation equation describes the correspondence between measured values ​​and state variables.

[0093] The system uses the lateral offset measured by Time-of-Flight (ToF) to construct the constraint equation: h(X(k))=y-[(LR) / 2]=0, where y is the lateral position component in the state vector. The state equation and observation equation are linearized at the current state estimate to obtain the Jacobian matrices F=∂f / ∂X and H=∂h / ∂X. The prediction equation of the extended Kalman filter is constructed as: X̂(k|k-1)=f(X̂(k-1|k-1), u(k)), P(k|k-1)=F(k)P(k-1|k-1)F(k)ᵀ+Q(k), where P is the state estimate covariance matrix and Q is the process noise covariance matrix.

[0094] S214. The state vector is iteratively updated using the extended Kalman filter model to obtain the fused localization result.

[0095] Among them, iterative update refers to the process of repeatedly predicting and correcting; the fused positioning results include three-dimensional position coordinates [x, y, z], three-dimensional attitude angles [roll, pitch, yaw] and their uncertainties; state vector iteration represents the gradual optimization of the state estimate during the filtering process.

[0096] The update process of the extended Kalman filter is performed in the following steps: First, calculate the Kalman gain K(k) = P(k|k-1)H(k)ᵀ[H(k)P(k|k-1)H(k)ᵀ+R(k)]⁻¹; then update the state estimate X̂(k|k) = X̂(k|k-1)+K(k)[Z(k)-h(X̂(k|k-1))]; finally, update the covariance matrix P(k|k) = [IK(k)H(k)]P(k|k-1). After iterative convergence, the following results are extracted from the final state vector X̂(k|k): (1) Three-dimensional position coordinates [x, y, z], representing the absolute position of the UAV in the world coordinate system; (2) Three-dimensional attitude angles [roll, pitch, yaw], describing the attitude of the UAV relative to the world coordinate system; (3) Three-dimensional velocity [vx, vy, vz], representing the motion rate in each direction; (4) Standard deviation of position estimation [σx, σy, σz], obtained from the square root of the diagonal elements of the covariance matrix, representing the uncertainty of the positioning result. The final output fused positioning result has the following performance indicators: horizontal position accuracy better than 5cm, vertical position accuracy better than 3cm, attitude angle accuracy better than 1 degree, and position update frequency not less than 50Hz.

[0097] In some embodiments, the method further includes: Acquire distance measurement data from a Time-of-Flight (ToF) sensor in the horizontal plane to determine the width and spacing of aisle walkways in a repetitive texture scene.

[0098] Among them, the distance measurement data in the horizontal plane refers to the distance value obtained by the ToF sensor on a plane parallel to the ground; the shelf aisle width refers to the interval distance between adjacent shelves; and the shelf spacing refers to the center distance between parallel shelf aisles.

[0099] The system controls the ToF sensor to perform distance measurements in the horizontal plane at a scanning angle of 60°, with a sampling interval of 1°, obtaining a distance data sequence {d1, d2, ..., d...}. 60 The aisle width *w* is obtained by detecting abrupt changes in distance data and identifying shelf edge locations. The distance between adjacent edge points is then calculated. The width values ​​of multiple aisles are statistically averaged to obtain the standard aisle width *W*. The shelf spacing *S* is calculated by determining the distance between the centerlines of adjacent aisles; typically, *S* equals the shelf depth plus the aisle width *W*.

[0100] Shelf edge features are extracted from image sequence data, and the extension direction of the shelf aisle is determined by combining distance measurement data from a time-of-flight (ToF) sensor.

[0101] Among them, the edge features of the shelf include geometric features such as straight line segments and corner points; the extension direction refers to the main axis direction of the shelf aisle.

[0102] The system performs edge detection and line extraction on image sequences, using Hough transform to identify parallel shelf edge segments. Simultaneously, it converts the Time-of-Flight (ToF) distance data into point clouds and fits the shelf plane using the RANSAC method. Combining the edge segment directions in the image with the point cloud plane normal vector, it calculates the extension direction vector v = [vx, vy, vz] of the shelf aisle. This direction vector is used for subsequent path planning and motion control.

[0103] When the Time-of-Flight (ToF) ranging sensor detects an obstacle ahead, it calculates the position of adjacent aisles based on the width and spacing of the aisle.

[0104] When the ToF sensor detects an obstacle less than a safe threshold (e.g., 2 meters) away, the system calculates the position coordinates of adjacent aisles based on the known shelf spacing S. Specifically, the center position of the adjacent aisle on the left is S meters to the left of the current position, and the center position of the adjacent aisle on the right is S meters to the right of the current position. The system evaluates the accessibility of the two adjacent aisles and selects the aisle with the least obstruction and the smallest deviation from the target direction as the turning target.

[0105] The drone is controlled to turn into adjacent aisles based on visual cumulative displacement, and its relative position to the shelf is maintained by utilizing the edge features of the shelf during the turning process.

[0106] The system continuously tracks the drone's position changes using visual odometry, generating a smooth turning trajectory from the current lane to the target lane. The trajectory uses a third-order Bézier curve to ensure speed continuity during the turn. During the turn, the system calculates the drone's lateral offset and heading deviation relative to the shelf by extracting shelf edge segments from the image in real time. A PID controller is used to adjust the drone's attitude, maintaining the desired relative position with the shelf and avoiding collisions.

[0107] Record the location information of the obstacle in the shelf aisle into the pre-stored map.

[0108] The system records the location information of detected obstacles in a pre-stored map. Specific recorded information includes: the index number of the passage where the obstacle is located, the distance of the obstacle from the passage entrance, the obstacle type identifier, and the detection timestamp. This information is stored in a structured format to facilitate passage reachability analysis during subsequent path planning. Map updates employ an incremental strategy, where new observation information overwrites older data at the same location.

[0109] Based on distance data measured by a Time-of-Flight (ToF) sensor, the width difference between two adjacent shelf aisles is calculated. When the width difference exceeds a preset width threshold, the shelf with the smaller width among the two adjacent shelf aisles is determined to be temporarily occupied. Spatial distribution features of obstacles are extracted from image sequence data and distance measurement data from the ToF sensor. When a change in spatial distribution features is detected, the changed area is marked as a dynamic obstacle avoidance area. During the execution of the obstacle avoidance trajectory, shelf aisles that are not marked as temporarily occupied or dynamic obstacle avoidance areas are selected as alternative obstacle avoidance paths.

[0110] The system first measures the actual widths W1 and W2 of adjacent shelf aisles using a ToF sensor, calculating the width difference ΔW = |W1 - W2|. When ΔW exceeds a preset threshold (e.g., 0.5 meters), the aisle with the smaller width is marked as temporarily occupied, and its location index, current width value, and detection timestamp are recorded. Simultaneously, the system performs target detection and tracking on the image sequence, extracting two-dimensional features of moving targets and converting the ToF distance data into a three-dimensional point cloud. Cluster analysis is used to identify independent obstacles in space, extracting the spatial location [x, y, z] and occupied volume V of each obstacle. The motion state of the obstacles is calculated through feature matching and position comparison between consecutive frames. When an obstacle's position offset exceeds 0.5 meters or its volume change exceeds 20%, the area and its surrounding 1-meter radius are marked as a dynamic obstacle avoidance zone, and the zone boundary coordinates and detection time are recorded. When obstacle avoidance is required, the system first acquires all adjacent shelf aisles reachable from the current location, excluding those marked as temporarily occupied and those intersecting with the dynamic obstacle avoidance zone through status checks. Among the remaining available channels, the optimal obstacle avoidance path is selected based on priority rules: closest to the target point, largest channel width, and smallest angle with the current direction of movement. After selecting the obstacle avoidance path, the system generates a trajectory that considers a 0.5-meter safety margin. During obstacle avoidance, the system continuously monitors environmental changes and updates the path planning in a timely manner when new dynamic obstacles are detected. This obstacle avoidance strategy based on multi-sensor fusion can effectively identify and respond to static occupancy and dynamic obstacle situations in the shelf environment, ensuring the safe operation of the drone.

[0111] In some embodiments, the method further includes: The angular velocity component is extracted from the inertial motion data. When the rate of change of the angular velocity component exceeds a preset angular velocity threshold, it is recognized that the drone is about to enter a corner scene. Within a preset time window before entering the corner scene, the fusion weight of visual positioning is adjusted from the first weight value to the second weight value through a linear gradient function. At the same time, the fusion weight of inertial positioning is adjusted from the third weight value to the fourth weight value, and the fusion weight of ToF-based geometric positioning is adjusted from the fifth weight value to the sixth weight value. The second weight value is less than the first weight value, the fourth weight value is greater than the third weight value, and the sixth weight value is greater than the fifth weight value. The scene type identifier and corridor width change sequence of the corresponding position of the grid cell are recorded in the grid cell of the pre-stored map. The scene type identifier includes the repeating texture scene identifier and the corner scene identifier. During flight, the scene type identifier of the scene ahead is obtained by querying the grid cell within a preset distance range based on the current position and flight direction. The sensor fusion weight adjustment of the corresponding scene is performed in advance based on the scene type identifier of the scene ahead.

[0112] The system first extracts the three-axis angular velocities [ωx, ωy, ωz] from the IMU data and calculates their rate of change dω / dt. When the rate of change of the angular velocity of any axis exceeds a preset threshold (e.g., 45° / s²), the system determines that the drone is about to enter a corner scene. After corner scene recognition is triggered, the system performs dynamic weight adjustment within a 1-second time window before entering the corner: the visual positioning weight is linearly reduced from 0.5 (first weight value) to 0.2 (second weight value) because image feature matching is prone to failure at corners; the inertial positioning weight is increased from 0.3 (third weight value) to 0.5 (fourth weight value) to provide continuous motion estimation; and the ToF-based geometric positioning weight is increased from 0.2 (fifth weight value) to 0.3 (sixth weight value) to enhance the dependence on wall structure features. The weight adjustment is implemented using a linear interpolation function w(t) = w1 + (w2 - w1) × t / T, where t is the current time and T is the time window length. Meanwhile, the system records environmental information in the pre-stored map using a grid-based storage structure: each grid cell (e.g., 0.5m × 0.5m) contains a scene type identifier (1 indicates a repeating texture scene, 2 indicates a corner scene) and a sequence of corridor width changes collected at that location {w1, w2, ..., w...}. n During flight, the system calculates the grid cell indices within a 5-meter radius ahead based on the current position [x, y] and heading angle θ, and queries the scene type identifiers of these grid cells. When a specific scene type is detected ahead, the corresponding weight adjustment strategy is activated in advance to achieve scene adaptation of the sensor fusion scheme. This scene prediction-based weight adjustment mechanism significantly improves the robustness and accuracy of multi-sensor fusion localization in complex environments.

[0113] The following describes the indoor positioning system for a micro-UAV in the embodiments of this invention from the perspective of hardware processing. Please refer to [link / reference needed]. Figure 3 This is a schematic diagram of the physical device structure of a micro unmanned aerial vehicle indoor positioning system in the embodiments of this application.

[0114] It should be noted that, Figure 3 The structure of the micro-drone indoor positioning system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0115] like Figure 3 As shown, the micro-UAV indoor positioning system includes a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from storage section 308 into Random Access Memory (RAM), such as executing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 303. The CPU 301, ROM 302, and RAM 303 are interconnected via bus 304. Input / output (I / O) interfaces are also connected to bus 304.

[0116] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0117] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.

[0118] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0120] Specifically, the indoor positioning system for micro unmanned aerial vehicles (UAVs) in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the ultrasonic environmental perception-based indoor positioning method for micro unmanned aerial vehicles provided in the above embodiment.

[0121] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the micro-UAV indoor positioning system described in the above embodiments; or it may exist independently and not assembled into the micro-UAV indoor positioning system. The storage medium carries one or more computer programs that, when executed by a processor of the micro-UAV indoor positioning system, cause the micro-UAV indoor positioning system to implement the ultrasonic environmental perception micro-UAV indoor positioning method provided in the above embodiments.

[0122] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0123] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An indoor positioning method for a micro unmanned aerial vehicle (UAV) using ultrasonic environmental perception, characterized in that, The method, applied to an indoor positioning system for micro unmanned aerial vehicles, includes: It acquires image sequence data from a vision camera, inertial motion data from an inertial measurement unit (IMU), and distance measurement data from a time-of-flight (ToF) sensor. The image sequence data is transformed in the frequency domain to obtain periodic spectral peaks. The distance between the left and right side walls is measured using the time-of-flight ranging sensor (ToF). When the variation amplitude of the periodic spectral peaks and the distance between the left and right side walls is within a preset threshold range, the current scene is identified as a repeating texture scene. In the repetitive texture scene, the inter-frame displacement transformation between adjacent image frames in the image sequence data is extracted, and the inter-frame displacement transformation is accumulated along the flight direction to obtain the visual cumulative displacement. The Time-of-Flight (ToF) sensor continuously measures the distances to the left and right side walls at multiple measurement locations along the flight direction, and then arranges the side wall distances at multiple measurement locations into a sequence of corridor width changes. The corridor width change sequence is matched with the reference width change sequence in the pre-stored map. The longitudinal positioning position is determined according to the position of the matched reference width change sequence in the map. The lateral offset is calculated according to the distance between the left and right side walls. The visual cumulative displacement, the longitudinal positioning position, the lateral offset, and the inertial motion data are input into a filter and fused to output a fused positioning result.

2. The method according to claim 1, characterized in that, The step of performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks, using the time-of-flight (ToF) ranging sensor to measure the distances to the left and right side walls, and identifying the current scene as a repeating texture scene when the variation amplitudes of the periodic spectral peaks and the distances to the left and right side walls are within a preset threshold range, specifically includes: A preset number of image frames are selected from the image sequence data to form an image analysis window. Feature points are extracted from each image frame in the image analysis window, and the number of feature point matches between adjacent image frames is calculated. When the number of matching feature points is greater than the first preset threshold, a two-dimensional fast Fourier transform is performed on the image frame in the image analysis window. Horizontal and vertical spectral components are extracted in the transformed frequency domain space, and the peak value of the spectral components is calculated as the periodic spectral peak value. The distance measurement sequence of the left and right side walls within a preset time period is obtained from the time-of-flight ranging sensor (ToF), and the root mean square error of the distance measurement sequence of the left and right side walls is calculated as the change range of the distance between the left and right side walls. When the peak value of the periodic spectrum is greater than the second preset threshold and the change in the distance between the left and right side walls is less than the third preset threshold, the current scene is determined to be the repeating texture scene.

3. The method according to claim 1, characterized in that, The step of performing sequence feature matching between the corridor width change sequence and a reference width change sequence in a pre-stored map, and determining the longitudinal positioning position based on the position of the matched reference width change sequence on the map, specifically includes: The corridor width change sequence is compared with the reference width change sequence using a sliding window, and the sequence correlation coefficient is calculated at each sliding window position; Obtain the position of the sliding window with the highest correlation coefficient, and take the starting position of the reference width change sequence corresponding to the sliding window position as the vertical positioning position; The distance from the current position to the central axis of the corridor is calculated based on the distance between the left and right side walls, and this distance is used as the lateral offset.

4. The method according to claim 1, characterized in that, The step of fusing the visual cumulative displacement, the longitudinal positioning position, the lateral offset, and the inertial motion data into the input filter and outputting the fused positioning result specifically includes: Acceleration and angular velocity components are extracted from the inertial motion data. Zero bias correction and integration are performed on the acceleration and angular velocity components to obtain the predicted inertial position value. A state vector is constructed, and a position observation equation is constructed using the inertial position prediction value, the visual cumulative displacement, and the longitudinal positioning position. The state vector includes position, velocity, and attitude components. A lateral constraint equation is constructed using the lateral offset, and an extended Kalman filter model is constructed based on the state vector, the position observation equation, and the lateral constraint equation. The state vector is iteratively updated using the extended Kalman filter model to obtain the fused localization result.

5. The method according to claim 1, characterized in that, After fusing the visual cumulative displacement, the longitudinal positioning position, the lateral offset, and the inertial motion data into the input filter and outputting the fused positioning result, the method further includes: Acquire distance measurement data of the Time-of-Flight (ToF) sensor in the horizontal plane to determine the width and spacing of the shelf aisles in the repeating texture scene; The shelf edge features are extracted from the image sequence data, and the extension direction of the shelf aisle is determined by combining the distance measurement data from the time-of-flight (ToF) sensor. When the Time-of-Flight (ToF) ranging sensor detects an obstacle ahead, it calculates the position of adjacent aisles based on the width and spacing of the shelf aisles. Based on the visual cumulative displacement control, the drone turns to the adjacent channel, and during the turning process, it uses the edge features of the shelf to maintain its relative position with the shelf; Record the location information of the obstacle in the shelf aisle into the pre-stored map.

6. The method according to claim 5, characterized in that, After the step of recording the location information of the obstacle in the shelf aisle into the pre-stored map, the method further includes: Based on the distance data measured by the Time-of-Flight (ToF) sensor, the width difference between two adjacent shelf aisles is calculated. When the width difference is greater than a preset width threshold, the shelf with the smaller width in the two adjacent shelf aisles is determined to be temporarily occupying the aisle; Spatial distribution features of obstacles are extracted from the image sequence data and the distance measurement data from the time-of-flight ranging sensor (ToF). When a change in the spatial distribution characteristics is detected, the changed area is marked as a dynamic obstacle avoidance area; During the execution of the obstacle avoidance trajectory, the shelf passage that is not marked as a temporarily occupied passage or a dynamic obstacle avoidance area is selected as the alternative obstacle avoidance path.

7. The method according to claim 1, characterized in that, After fusing the visual cumulative displacement, the longitudinal positioning position, the lateral offset, and the inertial motion data into the input filter and outputting the fused positioning result, the method further includes: Extract the angular velocity component from the inertial motion data, and when the rate of change of the angular velocity component exceeds a preset angular velocity threshold, identify that the drone is about to enter a cornering scene; Within a preset time window before entering the corner scene, the fusion weight of visual positioning is adjusted from a first weight value to a second weight value through a linear gradient function, while the fusion weight of inertial positioning is adjusted from a third weight value to a fourth weight value, and the fusion weight of geometric positioning based on ToF is adjusted from a fifth weight value to a sixth weight value. The second weight value is less than the first weight value, the fourth weight value is greater than the third weight value, and the sixth weight value is greater than the fifth weight value. In the grid cells of the pre-stored map, the scene type identifier and corridor width change sequence of the corresponding position of the grid cell are recorded. The scene type identifier includes repeating texture scene identifier and corner scene identifier. During flight, the system queries the grid cells within a preset distance range ahead based on the current position and flight direction to obtain the scene type identifier ahead, and performs sensor fusion weight adjustment for the corresponding scene in advance based on the scene type identifier ahead.

8. A micro unmanned aerial vehicle (UAV) indoor positioning system, characterized in that, The micro unmanned aerial vehicle (UAV) indoor positioning system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the micro unmanned aerial vehicle (UAV) indoor positioning system to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the micro unmanned aerial vehicle (UAV) indoor positioning system, the micro UAV indoor positioning system performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the micro unmanned aerial vehicle (UAV) indoor positioning system, the micro unmanned aerial vehicle (UAV) indoor positioning system performs the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle vision-inertia fusion indoor positioning method

    CN111024066A

  • AGV natural scene positioning method and system based on 3D laser radar

    CN120085317A

  • Unmanned aerial vehicle and ground sensing cooperative positioning method and system

    CN120668110A

  • Visual positioning method and device for tunnel scene and medium

    CN121033154A

  • Vehicle, vehicle positioning method and apparatus, device, and computer-readable storage medium

    WO2023065342A1