An indoor positioning method and system for a micro unmanned aerial vehicle based on ultrasonic environment perception
By using ultrasonic environmental perception methods, combined with visual and inertial measurement units (IMUs) and time-of-flight ranging (ToF) sensors, the positioning error of micro-UAVs in repetitive textured indoor scenes was identified and corrected, achieving stable indoor positioning and obstacle avoidance functions and improving the autonomous navigation capabilities of UAVs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-10
AI Technical Summary
In indoor environments, the visual positioning of micro drones is easily affected by repetitive texture scenes, leading to cumulative errors and positioning drift, which poses safety hazards.
An ultrasonic environmental perception method is employed, combining a visual camera, an inertial measurement unit (IMU), and a time-of-flight (ToF) ranging sensor. Through frequency domain analysis and multi-sensor data fusion, repetitive texture scenes are identified and their localization is corrected. Specific steps include: acquiring image sequence data and inertial motion data; performing frequency domain transformation to extract periodic spectral peaks; using ToF to measure the distance to the side walls; matching the corridor width variation sequence with a pre-stored map; and constructing an extended Kalman filter model to fuse the localization results.
It enables intelligent recognition and localization correction of repetitive texture scenes, improving the positioning accuracy and safety of micro drones in indoor environments and ensuring stable operation of drones in complex environments.
Smart Images

Figure CN121498718B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of non-electric variable control or regulation system, and particularly relates to an indoor positioning method and system of a micro unmanned aerial vehicle based on ultrasonic environment perception. BACKGROUND
[0002] With the rapid development of unmanned aerial vehicle technology, micro unmanned aerial vehicles are increasingly widely applied in indoor scenarios, and accurate indoor positioning capability is a key foundation for micro unmanned aerial vehicles to complete indoor inspection, material distribution and other tasks. In an indoor environment, GPS signals cannot be obtained, and other positioning means need to be relied on to realize autonomous navigation and positioning of the micro unmanned aerial vehicle.
[0003] At present, indoor unmanned aerial vehicle positioning mainly adopts a visual positioning scheme, image sequences are collected through a visual camera, image feature points are extracted and the motion of the feature points is tracked to estimate the position change of the unmanned aerial vehicle. In order to further improve the positioning effect, an inertial measurement unit (IMU) is introduced to obtain the attitude and motion data of the unmanned aerial vehicle, and the visual positioning result is fused.
[0004] However, in indoor corridor scenarios and the like, the wall surface often has a repetitive texture pattern, such as regularly arranged bricks or decorative stripes, and the repetitive texture can cause a large number of mis-matches in visual feature point extraction and tracking, so that the position estimation based on vision produces cumulative errors, and finally causes the unmanned aerial vehicle to deviate from the predetermined route and even collide. SUMMARY
[0005] The present application provides an indoor positioning method and system of a micro unmanned aerial vehicle based on ultrasonic environment perception, which is used to improve the accuracy of unmanned aerial vehicle positioning.
[0006] In a first aspect, the present application provides an indoor positioning method for a micro unmanned aerial vehicle based on ultrasonic environment perception. The method is applied to an indoor positioning system for a micro unmanned aerial vehicle, and comprises: collecting image sequence data of a vision camera, inertial motion data of an inertial measurement unit (IMU), and distance measurement data of a time-of-flight (ToF) ranging sensor; performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks, and measuring the distances to left and right side walls using the ToF ranging sensor; when the change amplitudes of the periodic spectral peaks and the distances to the left and right side walls are within a preset threshold range, identifying the current scene as a repetitive texture scene; in the repetitive texture scene, extracting an inter-frame displacement transformation amount between adjacent image frames in the image sequence data, and accumulating the inter-frame displacement transformation amount along a flight direction to obtain a visual accumulated displacement; continuously measuring the distances to the left and right side walls at a plurality of measurement positions along the flight direction using the ToF ranging sensor, and sequentially grouping the side wall distances at the plurality of measurement positions to obtain a corridor width change sequence; performing sequence feature matching on the corridor width change sequence and a reference width change sequence in a pre-stored map, determining a longitudinal positioning position according to a position of the matched reference width change sequence in the map, and calculating a lateral offset amount according to the distances to the left and right side walls; and inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset amount, and the inertial motion data into a filter for fusion, and outputting a fusion positioning result.
[0007] In the above embodiment, the system realizes intelligent identification and positioning correction capability for repetitive texture scenes. By fusing multi-source data of the vision camera, the IMU, and the ToF sensor, the system can identify and extract features of the repetitive texture scene in real time, calculate and correct the accumulated deviation of the visual positioning. At the same time, based on the corridor width change sequence measured by the ToF sensor, the system performs position matching, and cooperates with the motion information provided by the IMU to establish a reliable positioning solution model, thereby solving the problem of positioning drift of the traditional visual positioning in the repetitive texture environment.
[0008] In some embodiments of the first aspect, the step of performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks and measuring the left and right wall distances using the time-of-flight ranging sensor ToF, specifically comprises: selecting a preset number of image frames from the image sequence data to form an image analysis window, extracting feature points from each image frame in the image analysis window, and calculating the number of feature point matches between adjacent image frames; when the number of feature point matches is greater than a first preset threshold, performing two-dimensional fast Fourier transform on the image frames in the image analysis window, extracting spectral components in the horizontal and vertical directions in the transformed frequency domain space, and calculating the peak values of the spectral components as the periodic spectral peaks; obtaining a left and right wall distance measurement sequence within a preset time length from the time-of-flight ranging sensor ToF, and calculating the mean square deviation of the left and right wall distance measurement sequence as the variation amplitude of the left and right wall distances; when the periodic spectral peaks are greater than a second preset threshold and the variation amplitude of the left and right wall distances is less than a third preset threshold, determining that the current scene is a repetitive texture scene.
[0009] In the above embodiments, the system constructs a fast and accurate repetitive texture scene recognition mechanism. Combined with frequency domain analysis of image sequence and ToF distance measurement data, the system establishes a scene recognition model based on multi-feature fusion. By extracting the periodic spectral features of the image and the wall distance variation features, the system can accurately determine whether the current scene has repetitive texture. At the same time, an adaptive threshold judgment method is used to improve the robustness of scene recognition and provide a reliable basis for subsequent positioning strategy adjustment.
[0010] In some embodiments of the first aspect, the step of performing sequence feature matching between the corridor width variation sequence and the reference width variation sequence in the pre-stored map, and determining the longitudinal positioning position according to the position of the matched reference width variation sequence in the map, specifically comprises: comparing the corridor width variation sequence with the reference width variation sequence using a sliding window, calculating the sequence correlation coefficient of each sliding window position; obtaining the sliding window position with the maximum correlation coefficient, and taking the starting position of the reference width variation sequence corresponding to the sliding window position as the longitudinal positioning position; calculating the distance from the current position to the corridor center axis according to the left and right wall distances, and taking the distance as the lateral offset.
[0011] In the above embodiments, the system realizes accurate positioning based on corridor width variation. Using the continuous ranging of the left and right walls of the corridor by the ToF sensor, the system obtains the corridor width variation sequence reflecting the environmental features. By performing sliding window matching with the reference sequence in the pre-stored map, the system calculates the sequence correlation to determine the accurate position of the UAV. Combined with the calculation of the lateral offset, the system establishes a complete position estimation model to provide accurate position feedback for navigation control.
[0012] In some embodiments of the first aspect, the step of inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset and the inertial motion data into the filter to fuse and output a fusion positioning result specifically comprises: extracting acceleration components and angular velocity components from the inertial motion data, performing zero offset correction and integral operation on the acceleration components and the angular velocity components to obtain an inertial position prediction value; constructing a state vector including position, velocity and attitude components, constructing a position observation equation with the inertial position prediction value, the visual accumulated displacement and the longitudinal positioning position; constructing a lateral constraint equation with the lateral offset, and constructing an extended Kalman filter model according to the state vector, the position observation equation and the lateral constraint equation; and performing iterative update on the state vector by using the extended Kalman filter model to obtain the fusion positioning result.
[0013] In the above embodiments, the system realizes deep fusion of multi-sensor data. Based on the extended Kalman filter framework, the system constructs a state estimation model including position, velocity and attitude. Through the optimized fusion of visual accumulated displacement, IMU integral position and ToF geometric positioning result, a compensation mechanism for the measurement error of each sensor is established, and the accuracy and reliability of the positioning result are improved.
[0014] In some embodiments of the first aspect, after the step of inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset and the inertial motion data into the filter to fuse and output a fusion positioning result, the method further comprises: acquiring distance measurement data of a time-of-flight ranging sensor ToF in a horizontal plane, judging the width and spacing of the shelf channel in the repetitive texture scene; extracting shelf edge features from the image sequence data, and determining the extension direction of the shelf channel in combination with the distance measurement data of the time-of-flight ranging sensor ToF; when the time-of-flight ranging sensor ToF detects a front obstacle, calculating the position of the adjacent channel according to the width and spacing of the shelf channel; controlling the unmanned aerial vehicle to turn to the adjacent channel based on the visual accumulated displacement, and maintaining the relative position with the shelf during the turning process by using the shelf edge features; and recording the position information of the shelf channel where the obstacle is located into a pre-stored map.
[0015] In the above embodiments, the system constructs a complete intelligent obstacle avoidance function. Based on ToF horizontal plane scanning and visual feature recognition, the system realizes stereoscopic perception of environmental obstacles. By analyzing the geometric features and spatial distribution of the shelf channel, the system can calculate alternative obstacle avoidance paths, and realize smooth turning and obstacle avoidance in combination with visual feature tracking, thereby ensuring the safe operation ability of the unmanned aerial vehicle in complex environments.
[0016] In some embodiments of the first aspect, after the step of recording the position information of the shelf channel where the obstacle is located into the pre-stored map, the method further comprises: calculating the width difference of the two adjacent shelf channels based on the distance data measured by the time-of-flight ranging sensor ToF respectively; when the width difference is greater than a preset width threshold, determining that the shelf channel with smaller width among the two adjacent shelf channels is a temporary occupied channel; extracting the spatial distribution features of the obstacle from the image sequence data and the distance measurement data of the time-of-flight ranging sensor ToF; when a change in the spatial distribution features is detected, marking the changed area as a dynamic obstacle avoidance area; during the execution of the obstacle avoidance trajectory, selecting the shelf channel that is not marked as a temporary occupied channel and a dynamic obstacle avoidance area as a candidate obstacle avoidance path.
[0017] In the above embodiments, the system establishes a perception and processing mechanism for dynamic environments. By analyzing the width difference of adjacent shelf channels and the spatiotemporal distribution features of obstacles, the system can identify temporary occupied channels and dynamic change areas. Based on multi-level environmental feature analysis, the system realizes real-time updating of obstacle avoidance strategies, enhancing the adaptability and safety of drones in dynamic environments.
[0018] In some embodiments of the first aspect, after the step of inputting the visual accumulated displacement, longitudinal positioning position, lateral offset, and inertial motion data into the filter for fusion and outputting a fusion positioning result, the method further comprises: extracting the angular velocity component in the inertial motion data, and recognizing that the drone is about to enter a corner scene when the change rate of the angular velocity component exceeds a preset angular velocity threshold; within a preset time window before entering the corner scene, adjusting the fusion weight of visual positioning from a first weight value to a second weight value through a linear gradient function, adjusting the fusion weight of inertial positioning from a third weight value to a fourth weight value, and adjusting the fusion weight of ToF-based geometric positioning from a fifth weight value to a sixth weight value, the second weight value being less than the first weight value, the fourth weight value being greater than the third weight value, and the sixth weight value being greater than the fifth weight value; recording the scene type identifier and the corridor width change sequence corresponding to the position of the grid cell in the grid cell of the pre-stored map, the scene type identifier including a repetitive texture scene identifier and a corner scene identifier; during flight, querying the grid cell within a preset distance range in front to obtain a front scene type identifier according to the current position and flight direction, and performing sensor fusion weight adjustment for the corresponding scene in advance according to the front scene type identifier.
[0019] In the above embodiments, the system implements a sensor fusion mechanism for scene perception. By analyzing the IMU angular velocity data, the system can identify corner scenes in advance and adjust the fusion weights of different sensors through a linear gradient function. Combined with the scene type information in the pre-stored map, the system realizes predictive adjustment of the sensor fusion strategy and establishes a complete scene adaptive positioning framework.
[0020] In a second aspect, the embodiments of the present application provide a micro unmanned aerial vehicle indoor positioning system, comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is configured to store computer program codes, the computer program codes comprise computer instructions, the one or more processors invoke the computer instructions to enable the micro unmanned aerial vehicle indoor positioning system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0021] In a third aspect, the embodiments of the present application provide a computer program product comprising instructions which, when executed on a micro unmanned aerial vehicle indoor positioning system, enable the micro unmanned aerial vehicle indoor positioning system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0022] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium comprising instructions which, when executed on a micro unmanned aerial vehicle indoor positioning system, enable the micro unmanned aerial vehicle indoor positioning system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0023] It can be understood that the micro unmanned aerial vehicle indoor positioning system provided by the second aspect, the computer program product provided by the third aspect and the computer storage medium provided by the fourth aspect are all used to execute the method provided by the embodiments of the present application. Therefore, the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding method, which will not be repeated here.
[0024] The one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0025] 1. Since the multi-sensor fusion indoor positioning method is adopted, the complementary advantages of vision, IMU and ToF sensors can be fully utilized to realize intelligent recognition and positioning correction of repetitive texture scenes. The position matching is performed through the corridor width change sequence measured by the ToF sensor, and a positioning correction mechanism based on environmental geometric features is established. At the same time, a multi-source data fusion model is constructed by combining the motion information provided by the IMU and the cumulative displacement of vision, which effectively solves the problem that single visual positioning is easy to produce cumulative error in repetitive texture environment in the prior art, and further realizes stable and reliable indoor positioning function, and provides accurate position information support for autonomous navigation of unmanned aerial vehicles.
[0026] 2、Due to the adoption of the scene recognition method based on frequency domain analysis and ToF distance measurement, periodic spectral features can be extracted from the image sequence, and combined with the wall distance variation features measured by ToF, a multi-dimensional scene feature analysis model is constructed. By setting multiple preset threshold values for feature matching and judgment, a complete repetitive texture scene recognition process is established. The system can evaluate the current environmental features in real time and adjust the positioning strategy in a timely manner, effectively solving the problem of improper positioning strategy selection caused by the difficulty in identifying complex environmental features in the prior art, and thus realizing the adaptive adjustment of environmental perception and positioning strategy, and ensuring the positioning performance of the unmanned aerial vehicle in different scenes.
[0027] 3、Due to the adoption of the position matching method based on the corridor width variation sequence, the geometric structure features of the environment can be fully utilized for positioning correction. By continuous ranging of the ToF sensor at multiple positions, a width variation sequence reflecting the local environmental features is obtained. A sliding window sequence matching algorithm is adopted to calculate the correlation with the reference sequence in the pre-stored map, realizing large-scale positioning based on environmental features. By analyzing the left and right wall distances to calculate the lateral offset, the problem of positioning drift in repetitive texture areas in the prior art is effectively solved, and thus complete position estimation considering longitudinal positioning and lateral constraint is realized. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is a flowchart of an indoor positioning method of a micro unmanned aerial vehicle in an ultrasonic environment perception embodiment of the present application;
[0029] Figure 2 is another flowchart of an indoor positioning method of a micro unmanned aerial vehicle in an ultrasonic environment perception embodiment of the present application;
[0030] Figure 3 is a schematic diagram of an entity device structure of an indoor positioning system of a micro unmanned aerial vehicle in an embodiment of the present application. DETAILED DESCRIPTION
[0031] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be limiting to the present application. As used in the specification of the present application, the singular expression "one", "a", "the", "said" and "this" are intended to include the plural expression, unless there is clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application means any or all possible combinations of the listed items.
[0032] Hereinafter, the terms "first", "second", "third", "fourth", "fifth", "sixth", "seventh" and "eighth" are used only for descriptive purposes and should not be construed as implying or suggesting relative importance or an implied indication of the number of indicated technical features. Thus, the features defined with "first", "second", "third", "fourth", "fifth", "sixth", "seventh" and "eighth" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0033] For ease of understanding, the application scenarios of the embodiments of the present application are introduced as follows.
[0034] For ease of understanding, the method provided by the present embodiment is described in the following flow. Please refer to Figure 1 , a flowchart of an indoor positioning method of a micro unmanned aerial vehicle in an ultrasonic environment in the embodiments of the present application.
[0035] S101, collect image sequence data of a vision camera, inertial motion data of an inertial measurement unit (IMU), and distance measurement data of a time-of-flight (ToF) ranging sensor.
[0036] Among them, the image sequence data represents a data stream composed of multiple frames of images continuously collected by the vision camera at a fixed frequency; the inertial measurement unit (IMU) refers to a sensor device for measuring the acceleration and angular velocity of an object; the time-of-flight (ToF) ranging sensor refers to a sensor device that measures distance by transmitting and receiving light signals.
[0037] This step is performed when the unmanned aerial vehicle takes off and starts to perform the indoor positioning task. Specifically, the system collects a sequence of RGB images through the vision camera at a frequency of not less than 30 frames per second, collects three-axis acceleration and three-axis angular velocity data output by the IMU at a frequency of more than 200 Hz, and collects distance measurement data output by the ToF sensor at a frequency of not less than 10 Hz, to realize the synchronous collection of multi-sensor data.
[0038] In some embodiments, the collection and synchronization of multi-sensor data can be realized in various ways: optionally, the sampling clocks of the three sensors are synchronized using a hardware trigger signal, and the sampled data is buffered into three FIFO queues, and data alignment and fusion are performed according to timestamp information; optionally, a software timer is used to trigger sensor sampling, and time interpolation and data synchronization are performed in the data processing thread. It can be understood that other sensor synchronization schemes can also be used to realize data collection.
[0039] S102, performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks, measuring the left and right side wall distances using the time-of-flight (ToF) ranging sensor, and when the change amplitudes of the periodic spectral peaks and the left and right side wall distances are within a preset threshold range, identifying the current scene as a repetitive texture scene.
[0040] Wherein, the frequency domain transform represents a mathematical transform method of converting a time domain signal to a frequency domain; the periodic spectral peak refers to an amplitude of a periodic characteristic exhibited by a signal in a frequency domain; the variation amplitude represents a degree of fluctuation of a signal; and the repetitive texture scene refers to a scene in which a regular repeating pattern exists in an environment.
[0041] This step is executed after data acquisition is completed. Specifically, the system first performs a fast Fourier transform on the acquired image sequence, analyzes the spectral characteristics of the image in the horizontal and vertical directions, and extracts the peak values of the main frequency components. Meanwhile, the ToF sensor is used to measure the distances of the left and right walls, and the statistical characteristics of the distance values are calculated. When the spectral peak and the distance variation characteristics both satisfy the preset conditions, it is determined that the current scene is a repetitive texture scene.
[0042] In some embodiments, scene recognition can be achieved in various ways: optionally, a sliding window method is used to process the image sequence, a two-dimensional FFT transform is performed on each window, the spectral peak values in the horizontal and vertical directions are extracted, and scene judgment is performed in combination with the variance characteristics of the ToF distance data; optionally, a deep learning method is used to train a scene classification model, the spectral characteristics and distance characteristics are used as inputs, and the scene type probability is output. It can be understood that other methods can also be used to achieve scene recognition.
[0043] S103、In the repetitive texture scene, the inter-frame displacement transform amount between adjacent image frames in the image sequence data is extracted, and the visual cumulative displacement is obtained by accumulating the inter-frame displacement transform amount along the flight direction.
[0044] Wherein, the inter-frame displacement transform amount represents the relative motion displacement between adjacent image frames; the visual cumulative displacement refers to the total displacement amount accumulated along the motion direction; and the flight direction refers to the forward direction of the UAV.
[0045] This step is executed after the repetitive texture scene is identified. Specifically, the system extracts feature points from adjacent frames in the image sequence and performs matching, calculates the inter-frame homography matrix, and decomposes to obtain the relative motion parameters. The inter-frame displacement is accumulated along the flight direction to obtain the visual-based cumulative displacement estimation value.
[0046] In some embodiments, visual displacement estimation can be achieved in various ways: optionally, a FAST corner detection algorithm is used to extract feature points, an optical flow method is used to track the motion of the feature points, and the inter-frame motion parameters are calculated; optionally, a feature point matching method is used to estimate the fundamental matrix, and the camera motion is decomposed. It can be understood that other visual odometry methods can also be used to achieve displacement estimation.
[0047] S104、The left and right side wall distances are continuously measured at multiple measurement positions along the flight direction by a time-of-flight ranging sensor ToF, and the side wall distances at the multiple measurement positions are sequentially combined to form a corridor width variation sequence.
[0048] wherein, the measurement positions represent the spatial sampling points where the ToF sensor measures the distance; the left and right wall distances are the actual distance values from the sensor to the left and right side walls of the corridor; the corridor width variation sequence represents an ordered data set of the corridor width at multiple positions sampled along the flight direction; the continuous measurement refers to the distance measurement process that keeps coherent in time and space.
[0049] This step is performed immediately after obtaining the sensor data. Specifically, the system controls the ToF sensor to sample at fixed spatial intervals (e.g., 0.5 meters) along the flight path, and measures the left and right wall distances at each sampling position respectively. The sum of the measured left and right distance values gives the corridor width at that position, and the width values of the consecutive N positions (e.g., N = 20) are arranged in time sequence to form a sequence. This sequence reflects the local variation characteristics of the corridor geometry, which can be used for subsequent position matching.
[0050] In some embodiments, the construction of the corridor width sequence can be achieved in various ways: alternatively, multiple ToF sensors can be used to measure distances in different directions simultaneously, and accurate wall distances can be obtained through coordinate transformation and data fusion; alternatively, IMU data can be combined to compensate for the attitude of the ToF measurement values, eliminating measurement errors caused by changes in flight attitude; alternatively, Kalman filtering can be used to smooth the original measurement data, improving the reliability of the sequence data. It can be understood that other methods can also be used to obtain the corridor width sequence.
[0051] It should be noted that the "corridor width variation sequence" described in this application not only refers to the gradual change of the macroscopic physical width of the corridor, but also includes detailed information of the local geometric features of the walls on both sides of the corridor. In actual indoor environments, even a macroscopically equal-width corridor (e.g., an office building corridor with a constant width of 2 meters), its walls are not absolutely flat surfaces, but are distributed with structures having significant geometric features.
[0052] Specifically, the features relied on by the feature matching include but are not limited to:
[0053] Incremental features (recesses): such as door hole recesses on both sides of the corridor, open doors, fire hydrant installation slots, elevator entrance, etc. These features are manifested as instantaneous stepwise increases in width (pulse wave peaks) in ToF measurement data;
[0054] Decreasing features (protrusions): such as load-bearing columns, wall decoration protrusions, exposed pipe wells, etc. These features are manifested as instantaneous decreases in width (pulse troughs) in ToF measurement data;
[0055] Texture features: reflection rate mutation points caused by changes in wall material (requires ToF sensor to support intensity echo).
[0056] Therefore, the "corridor width variation sequence" constructed in this embodiment is actually a set of "one-dimensional depth fingerprints" reflecting the environmental geometry. Even in a long straight and equal-width corridor, as long as there is any of the above-mentioned geometric features, the sequence is unique, so that the longitudinal position can be determined by feature matching.
[0057] S105, sequence feature matching is performed between the corridor width variation sequence and the reference width variation sequence in the pre-stored map, the longitudinal positioning position is determined according to the position of the matched reference width variation sequence in the map, and the lateral offset is calculated according to the distance between the left and right side walls.
[0058] In the formula, the pre-stored map represents a database containing environmental geometric information established in advance, the pre-stored map not only records the width of the corridor, but also marks the position coordinates and size information of the key geometric features (such as doors and columns); the reference width variation sequence refers to the standard corridor width variation pattern recorded in the map; the sequence feature matching refers to the matching degree between the two sequences; the longitudinal positioning position represents the position coordinates in the direction of the corridor; and the lateral offset represents the vertical distance to the center line of the corridor.
[0059] This step is performed after the corridor width variation sequence is obtained. Specifically, the system performs sliding window matching between the currently obtained width variation sequence and all reference sequences stored in the map database, and calculates the normalized cross-correlation coefficient. The starting position of the corresponding reference sequence in the map is selected as the current longitudinal position when the matching result has the maximum correlation coefficient and exceeds the threshold. At the same time, the lateral offset distance of the unmanned aerial vehicle relative to the center line of the corridor is calculated according to the distance difference between the left and right walls.
[0060] In some embodiments, the position matching and offset calculation can be realized in various ways: optionally, a dynamic time warping (DTW) algorithm is used for sequence matching to adapt to the nonlinear characteristics of the corridor width variation; optionally, a multi-scale sequence template is constructed to improve the positioning efficiency and robustness through hierarchical matching; optionally, the wall material and texture features are combined to assist sequence matching to improve the accuracy of position recognition. It can be understood that other methods can also be used to realize position matching.
[0061] S106, the visual cumulative displacement, the longitudinal positioning position, the lateral offset and the inertial motion data are input into a filter for fusion, and a fusion positioning result is output.
[0062] In the formula, the filter represents a state estimator for multi-source data fusion; the fusion positioning result refers to the optimal position estimation obtained by integrating various sensor information; and the inertial motion data includes acceleration and angular velocity measurement values.
[0063] This step is performed after obtaining various position estimates. Specifically, the system constructs an extended Kalman filter, taking the visual displacement, ToF-based position matching result and IMU integrated position as observations, and establishes state transition equations and observation equations. Through the iterative prediction and update process, the advantages of each sensor data are fused, and the cumulative errors of each are suppressed, and finally a reliable position estimate result is output.
[0064] In some embodiments, multisensor data fusion can be achieved in various ways: optionally, an adaptive weight method based on covariance matrix is adopted to dynamically adjust the fusion weight according to the reliability of each sensor data; optionally, a particle filter is introduced to process the nonlinear observation model, improving the positioning accuracy; optionally, the scene recognition result is combined to adjust the filter parameters, adapting to the positioning needs in different environments. It can be understood that other methods can also be used to achieve data fusion.
[0065] The method provided by the present embodiment is further described in more detail below. Please refer to Figure 2 , another flowchart of the indoor positioning method of the ultrasonic environment sensing micro unmanned aerial vehicle in the embodiment of the present application.
[0066] S201, collect image sequence data of a vision camera, inertial motion data of an inertial measurement unit (IMU) and distance measurement data of a time-of-flight (ToF) ranging sensor.
[0067] In some embodiments, this step is similar to step S101, which will not be described again here.
[0068] S202, select a preset number of image frames from the image sequence data to form an image analysis window, extract feature points from each image frame in the image analysis window, and calculate the number of feature point matches between adjacent image frames.
[0069] Wherein, the preset number represents a fixed number of image frames, usually 8-12 frames are selected to ensure the time continuity and calculation efficiency of the analysis; the image analysis window refers to a set of continuous image frames for feature extraction and motion analysis, reflecting the change characteristics of the scene in a short time; the feature point represents a significant corner or edge point in the image, which has the characteristic of obvious local area gray change; the number of feature point matches refers to the number of successfully paired feature points between two images, reflecting the similarity of image content and motion continuity.
[0070] The system selects 10 continuous images from the image sequence buffer as an analysis window in real time, and extracts feature points in each image using the FAST corner detection algorithm. The algorithm first selects candidate pixel points in the image, and compares the gray value difference between the center point and its surrounding 16 pixel points. When the gray value difference between the center point and its surrounding N (N>9) pixel points exceeds the threshold, the point is marked as a corner. For each detected corner, calculate its BRIEF binary descriptor, which is obtained by comparing the gray values of a number of pixel pairs randomly selected in the neighborhood of the corner. Calculate the Hamming distance of the feature point descriptors of adjacent two images, and use the nearest neighbor ratio rule (threshold set to 0.7) to screen matching pairs to obtain the initial matching result. Further use RANSAC algorithm to remove outliers to obtain the final effective matching point pair. For images with a resolution of 640x480, usually 300-500 feature points are extracted per frame, and 100-200 stable matching pairs can be obtained between adjacent frames. The system records the number of matching pairs, which is used for subsequent scene feature analysis.
[0071] In S203, when the number of feature point matches is greater than a first preset threshold, a two-dimensional fast Fourier transform is performed on the image frames in the image analysis window, and the spectral components in the horizontal and vertical directions are extracted in the transformed frequency domain space. The peak value of the spectral components is calculated as the periodic spectral peak value.
[0072] The two-dimensional fast Fourier transform is an efficient mathematical operation method for converting a two-dimensional image signal from the spatial domain to the frequency domain, which can reveal the periodic structural features of the image. The frequency domain space represents a two-dimensional coordinate system of signal frequency distribution, and the horizontal and vertical coordinates correspond to the horizontal and vertical spatial frequencies, respectively. The spectral component refers to the energy distribution of the signal at different frequencies, and its amplitude reflects the strength of the corresponding frequency component. The periodic spectral peak value represents a significant local maximum in the frequency spectrum, corresponding to the main periodic feature in the image.
[0073] When the number of feature point matches exceeds a set threshold (e.g., 150 pairs), the system performs the following processing on each image in the analysis window: first, the image is converted to a grayscale image and histogram equalization is performed to enhance the image contrast. Then, a two-dimensional FFT transform is performed on the pre-processed image to obtain a complex matrix in the frequency domain. The modulus of the complex matrix is calculated to obtain a power spectrum, and a logarithmic transform is performed to enhance the visual effect. One-dimensional spectrum curves are obtained by accumulating the power spectrum along the horizontal direction (u-axis) and the vertical direction (v-axis), respectively. A sliding window method (window size is 1 / 10 of the total length of the spectrum) is used to detect peaks in the spectrum curve, and local maximum points greater than twice the mean value in the window are recorded. The detected peaks are sorted by amplitude, and the largest peak is selected as the quantification index of the periodic feature. For example, for an image with regular brick wall texture, a significant spectrum peak will appear at the spatial frequency corresponding to the brick spacing, and the peak amplitude is usually 3-5 times the mean value of the background spectrum.
[0074] S204, obtaining a left and right side wall distance measurement sequence within a preset time length from a time-of-flight ranging sensor ToF, and calculating a mean square deviation of the left and right side wall distance measurement sequence as a change amplitude of the left and right side wall distance.
[0075] wherein the two-dimensional fast Fourier transform is an efficient mathematical operation method for converting a two-dimensional image signal from a spatial domain to a frequency domain, which can reveal the periodic structural features of the image; the frequency domain space represents a two-dimensional coordinate system of signal frequency distribution, the horizontal and vertical coordinates correspond to the spatial frequencies in the horizontal and vertical directions, respectively; the spectrum component refers to the energy distribution of the signal at different frequencies, and the amplitude reflects the intensity of the corresponding frequency component; the periodic spectrum peak represents a significant local maximum value in the spectrum, corresponding to the main periodic feature in the image.
[0076] When the number of feature point matches exceeds a set threshold (e.g., 150 pairs), the system performs the following processing on each image in the analysis window: first, the image is converted to a grayscale image and histogram equalization is performed to enhance the image contrast. Then, a two-dimensional FFT transform is performed on the pre-processed image to obtain a complex matrix in the frequency domain. The modulus of the complex matrix is calculated to obtain a power spectrum, and a logarithmic transform is performed to enhance the visual effect. One-dimensional spectrum curves are obtained by accumulating the power spectrum along the horizontal direction (u-axis) and the vertical direction (v-axis), respectively. A sliding window method (window size is 1 / 10 of the total length of the spectrum) is used to detect peaks in the spectrum curve, and local maximum points greater than twice the mean value in the window are recorded. The detected peaks are sorted by amplitude, and the largest peak is selected as the quantification index of the periodic feature. For example, for an image with regular brick wall texture, a significant spectrum peak will appear at the spatial frequency corresponding to the brick spacing, and the peak amplitude is usually 3-5 times the mean value of the background spectrum.
[0077] S205: When the periodic spectral peak is greater than the second preset threshold and the change range of the left and right wall distance is less than the third preset threshold, it is determined that the current scene is a repetitive texture scene.
[0078] wherein the preset time length represents a fixed data acquisition time window, usually 1-3 seconds to ensure the reliability of data statistics; the distance measurement sequence refers to a set of ToF sensor distance data continuously acquired within the sampling time; the mean square error represents a statistical index of the degree of data fluctuation, reflecting the stability of the wall distance.
[0079] The system controls the ToF sensor to measure the left and right wall distances simultaneously within a preset 2-second time window at a sampling frequency of 50 Hz. For each measurement, the sensor emits a modulated near-infrared light pulse, and the distance value is calculated by detecting the phase delay of the reflected light signal. Within 2 seconds, about 100 left wall distance sampling points {L1, L2,..., L 100} and right wall distance sampling points {R1, R2,..., R 100} are obtained respectively. The arithmetic mean values μL and μR are calculated for the two sequences respectively, and then the sum of the squares of the differences between each sampling point and the mean value is calculated, and the square root of the sampling point number is obtained to get the mean square errors σL and σR of the left and right wall distance sequences. The specific calculation formula is: σL = sqrt(Σ(Li-μL)² / 100) and σR = sqrt(Σ(Ri-μR)² / 100). The larger mean square error value max(σL, σR) is selected as the measure of the change range of the wall distance. When this value is less than the preset threshold (such as 5 cm), it indicates that the wall distance is relatively stable, which helps to determine that the current is in a regular corridor environment. At the same time, the system also records the absolute values of the left and right wall distances, which are used for subsequent position estimation and path planning.
[0080] S206: In the repetitive texture scene, the inter-frame displacement transformation amount between adjacent image frames in the image sequence data is extracted, and the visual cumulative displacement is obtained by accumulating the inter-frame displacement transformation amount along the flight direction.
[0081] In some embodiments, this step is similar to step S103, which will not be described here.
[0082] S207: Continuously measure the left and right wall distances along the flight direction at multiple measurement positions by the time-of-flight ranging sensor ToF, and sequentially form a corridor width change sequence by the wall distances at the multiple measurement positions.
[0083] In some embodiments, this step is similar to step S104, which will not be described here.
[0084] S208: Compare the corridor width change sequence with the reference width change sequence in a sliding window, and calculate the sequence correlation coefficient at each sliding window position.
[0085] S209. Obtain the position of the sliding window with the largest correlation coefficient, and take the starting position of the reference width change sequence corresponding to the sliding window position as the vertical positioning position.
[0086] S210. Calculate the distance from the current position to the central axis of the corridor based on the distance between the left and right side walls, and use the distance as the lateral offset. Calculate the lateral offset based on the distance between the left and right side walls.
[0087] Among them, the sliding window represents a fixed-length data segment that moves along the sequence and is used for local sequence alignment; the sequence correlation coefficient is a statistic that measures the similarity between two data sequences, with a value range of [-1, 1]; the longitudinal positioning position refers to the absolute position coordinates of the UAV in the direction of the corridor; the corridor centerline refers to the center line of the walls on both sides of the corridor; and the lateral offset represents the vertical distance of the UAV from the corridor centerline.
[0088] The system first sets the sliding window length W (e.g., 20 sampling points), and then processes the currently obtained corridor width variation sequence {w1, w2, ..., w...}. n} and the reference sequence {r1, r2, ..., r} in the pre-stored map m A sliding comparison is performed. For each window position i, the Pearson correlation coefficient of the sequence segments within the window is calculated: ρᵢ = cov(W_i, R_i) / [std(W_i)×std(R_i)], where W_i and R_i are the data segments of the two sequences at the current window position, cov represents the covariance, and std represents the standard deviation. All possible window positions are iterated to find the position i_max with the highest correlation coefficient. The starting coordinates of the reference sequence corresponding to the i_max position are taken as the longitudinal position of the UAV. Simultaneously, based on the currently measured distances L and R between the left and right walls, the lateral offset is calculated: offset = (LR) / 2. This value represents the distance the UAV deviates from the corridor centerline; a positive value indicates a deviation to the right, and a negative value indicates a deviation to the left. This yields the complete two-dimensional position of the UAV in the corridor coordinate system.
[0089] Furthermore, for the special scenario of "equal width corridors", the process of matching the corridor width change sequence with the reference width change sequence in the pre-stored map preferably adopts a local matching strategy based on salient features.
[0090] The specific steps are as follows: The acquired real-time width sequence is preprocessed, and the first derivative or gradient of the sequence is calculated to identify abrupt change points (i.e., geometric feature points) in the data. For example, when a drone flies over a doorway, the ToF measurement of the distance to the left or right will generate a rectangular wave signal exhibiting a "distance increase-hold-decrease" pattern. The system extracts the feature attributes of these abrupt change points, including: feature type (recess or protrusion), feature width (duration along the flight direction), feature depth (depth of the recess or height of the protrusion), and the spacing between adjacent features. The extracted real-time feature attribute set is then topologically matched with the reference feature attribute set recorded in a pre-stored map. For example, if the real-time sequence detects "a recess (doorway) with a width of 0.9 meters and a depth of 0.2 meters, and a protrusion (pillar) with a width of 0.5 meters detected 3 meters in front of it," the system will search the map database for areas with the same topological structure. In this way, even in extreme environments where visual textures are highly repetitive (such as white walls) and the overall width of the corridor is constant, the system can still accurately lock the longitudinal position of the drone by utilizing the distribution patterns of geometric features such as doorways and pillars, effectively solving the drift problem of traditional visual odometry in long corridor scenarios.
[0091] S211. Extract acceleration and angular velocity components from inertial motion data, perform zero-bias correction and integration on acceleration and angular velocity components, and obtain the predicted inertial position value.
[0092] Among them, the acceleration component refers to the three-axis linear acceleration measured by the IMU; the angular velocity component is the three-axis angular rate measured by the IMU; zero bias correction represents the process of eliminating the static error of the sensor; integration operation refers to the time integration of acceleration and angular velocity data.
[0093] The system extracts triaxial acceleration [ax, ay, az] and angular velocity [ωx, ωy, ωz] from the raw IMU data. First, zero-bias correction is performed: 5 seconds of data are collected while the system is stationary, and the average values of each axis are calculated as zero biases [bax, bay, baz] and [bωx, bωy, bωz]. These zero-bias values are subtracted from the real-time data. The corrected angular velocity is integrated in first order to obtain the attitude angle change, and the rotation matrix R from the computational volume coordinate system to the world coordinate system is updated using quaternions. The acceleration data is transformed to the world coordinate system using R, and after subtracting the gravitational acceleration g, a double integration is performed: the first integration yields the velocity v(t) = ∫(a(t) - g)dt, and the second integration yields the position p(t) = ∫v(t)dt. The integration is implemented using the trapezoidal rule, with the integration time step set to the IMU sampling period (e.g., 0.005 seconds). The resulting p(t) is the predicted position value based on the inertial data.
[0094] S212, construct a state vector including position, velocity and attitude components to construct a position observation equation with inertial position prediction, visual accumulated displacement and longitudinal positioning position.
[0095] wherein the state vector is a collection of variables describing the dynamic characteristics of the system; the position component contains three-dimensional space coordinates [x, y, z]; the velocity component represents the motion rate in three directions [vx, vy, vz]; the attitude component is represented by Euler angles [roll, pitch, yaw]; and the position observation equation describes the functional relationship between the measured value and the state variable.
[0096] The system constructs a 13-dimensional state vector X = [x, y, z, vx, vy, vz, roll, pitch, yaw, bax, bay, baz, bω], which contains position, velocity, attitude angle and IMU zero offset. Based on the state vector, a nonlinear system state equation is constructed: X(k+1) = f(X(k), u(k)) + w(k), wherein u(k) is the IMU input and w(k) is the process noise. The position observation equation adopts a multi-source fusion form: Z(k) = [Zimu(k); Zvis(k); Ztof(k)], which corresponds to inertial position prediction, visual accumulated displacement and longitudinal position measurement based on ToF respectively. The measurement noise covariance matrix R of the observation equation is set according to the measurement accuracy of each sensor.
[0097] S213, construct a lateral constraint equation with a lateral offset, and construct an extended Kalman filter model according to the state vector, the position observation equation and the lateral constraint equation.
[0098] wherein the lateral constraint equation represents the geometric constraint of the UAV in the lateral movement of the corridor; the extended Kalman filter model is a state estimator for processing nonlinear systems; and the position observation equation describes the corresponding relationship between the measured value and the state variable.
[0099] The system constructs a constraint equation with the lateral offset measured by ToF: h(X(k)) = y - [(L - R) / 2] = 0, wherein y is the lateral position component in the state vector. The state equation and the observation equation are linearized at the current state estimation value to obtain the Jacobian matrices F = ∂f / ∂X and H = ∂h / ∂X. The prediction equation of the extended Kalman filter is constructed: X̂(k|k-1) = f(X̂(k-1|k-1), u(k)), P(k|k-1) = F(k)P(k-1|k-1)F(k)ᵀ+Q(k), wherein P is the state estimation covariance matrix and Q is the process noise covariance matrix.
[0100] S214, iteratively update the state vector using the extended Kalman filter model to obtain a fusion positioning result.
[0101] wherein the iterative update refers to a process of repeatedly performing prediction and correction; the fusion positioning result includes three-dimensional position coordinates [x, y, z], three-dimensional attitude angles [roll, pitch, yaw] and their uncertainties; and the state vector iteration represents step-by-step optimization of state estimation values in the filtering process.
[0102] The update process of the extended Kalman filter is performed according to the following steps: first, calculate the Kalman gain K(k) = P(k|k-1)H(k)ᵀ[H(k)P(k|k-1)H(k)ᵀ+R(k)]⁻¹; then update the state estimation X̂(k|k) = X̂(k|k-1) + K(k)[Z(k) - h(X̂(k|k-1))]; and finally update the covariance matrix P(k|k) = [I - K(k)H(k)]P(k|k-1). After iterative convergence, the following results are extracted from the final state vector X̂(k|k): (1) three-dimensional position coordinates [x, y, z], representing the absolute position of the UAV in the world coordinate system; (2) three-dimensional attitude angles [roll, pitch, yaw], describing the attitude of the UAV relative to the world coordinate system; (3) three-dimensional velocity [vx, vy, vz], representing the motion rate in each direction; and (4) standard deviation of position estimation [σx, σy, σz], obtained by taking the square root of the diagonal elements of the covariance matrix, representing the uncertainty of the positioning result. The final output fusion positioning result has the following performance indicators: horizontal position accuracy better than 5 cm, vertical position accuracy better than 3 cm, attitude angle accuracy better than 1 degree, and position update frequency not less than 50 Hz.
[0103] In some embodiments, the method further comprises:
[0104] Obtaining distance measurement data of the time-of-flight ranging sensor (ToF) in the horizontal plane, and determining the width and spacing of the shelf aisle in the repetitive texture scene.
[0105] wherein the distance measurement data in the horizontal plane refers to the distance values obtained by the ToF sensor in a plane parallel to the ground; the shelf aisle width refers to the interval distance between adjacent shelves; and the shelf spacing represents the center distance between parallel shelf aisles.
[0106] The system controls the ToF sensor to perform distance measurement in the horizontal plane with a scanning angle of 60°, a sampling interval of 1°, and obtains a distance data sequence {d1, d2,..., d 60}. The shelf edge positions are identified by detecting the abrupt points of the distance data, and the distance between adjacent edge points is calculated to obtain the aisle width w. The width values of multiple aisles are statistically averaged to obtain the standard aisle width W. The distance between the center lines of adjacent aisles is calculated to obtain the shelf spacing S, which is usually equal to the shelf depth plus the aisle width W.
[0107] Extracting the shelf edge features from the image sequence data, combined with the distance measurement data of the time-of-flight ranging sensor ToF to determine the extension direction of the shelf aisle.
[0108] Wherein, the shelf edge features include straight line segments, corner points and other geometric features; the extension direction refers to the main axis direction of the shelf aisle.
[0109] The system performs edge detection and line extraction on the image sequence, and uses Hough transform to identify parallel shelf edge line segments. At the same time, the ToF distance data is converted into point cloud, and the shelf plane is fitted by RANSAC method. Combined with the direction of the edge line segment in the image and the normal vector of the point cloud plane, the extension direction vector v = [vx, vy, vz] of the shelf aisle is calculated. The direction vector is used for subsequent path planning and motion control.
[0110] When the time-of-flight ranging sensor ToF detects an obstacle in front, the positions of adjacent aisles are calculated according to the width and spacing of the shelf aisle.
[0111] When the ToF sensor detects that the distance to the obstacle in front is less than a safety threshold (such as 2 meters), the system calculates the position coordinates of the adjacent aisles according to the known shelf spacing S. Specifically, the center position of the left adjacent aisle is S meters to the left of the current position, and the center position of the right adjacent aisle is S meters to the right of the current position. The system evaluates the accessibility of the left and right adjacent aisles and selects the aisle with no obstacles and the smallest deviation from the target direction as the turning target.
[0112] Based on the visual cumulative displacement, the UAV is controlled to turn into the adjacent aisle, and the relative position with the shelf is maintained during the turning process.
[0113] The system continuously tracks the position change of the UAV based on the visual odometry, and generates a smooth turning trajectory from the current aisle to the target aisle. The trajectory adopts a third-order Bezier curve form to ensure the continuity of the turning process. During the turning process, the shelf edge line segments in the image are extracted in real time, and the lateral deviation and heading deviation of the UAV relative to the shelf are calculated. A PID controller is used to adjust the attitude of the UAV to maintain the desired relative position with the shelf and avoid collisions.
[0114] The position information of the obstacle in the shelf aisle is recorded in the pre-stored map.
[0115] The system records the position information of the detected obstacle in the pre-stored map. The specific information recorded includes: the index number of the aisle where the obstacle is located, the distance of the obstacle from the entrance of the aisle, the type identification of the obstacle, the detection timestamp, etc. These information is stored in a structured format, which is convenient for subsequent path planning and aisle accessibility analysis. The map update adopts an incremental strategy, and the new observation information will overwrite the old data at the same position.
[0116] The width difference of two adjacent shelf channels is calculated based on the distance data measured by the time-of-flight ranging sensor ToF, and when the width difference is greater than a preset width threshold, the shelf with smaller width in the two adjacent shelf channels is determined as a temporarily occupied channel. The spatial distribution features of the obstacles are extracted from the image sequence data and the distance measurement data of the time-of-flight ranging sensor ToF, and when a change in the spatial distribution features is detected, the changed area is marked as a dynamic obstacle avoidance area. In the process of executing the obstacle avoidance trajectory, the shelf channel that is not marked as the temporarily occupied channel and the dynamic obstacle avoidance area is selected as a candidate obstacle avoidance path.
[0117] The system first measures the actual widths W1 and W2 of the adjacent shelf channels by the ToF sensor, and calculates the width difference AW = |W1-W2|. When AW is greater than a preset threshold (such as 0.5 meters), the channel with smaller width is marked as a temporarily occupied channel, and the position index, current width value and detection timestamp of the channel are recorded. At the same time, the system performs target detection and tracking on the image sequence, extracts the two-dimensional features of the moving targets, and converts the ToF distance data into three-dimensional point cloud. Through clustering analysis, independent obstacles in space are identified, and the spatial position [x, y, z] and the occupied volume V of each obstacle are extracted. Through feature matching and position comparison between consecutive frames, the motion state of the obstacle is calculated. When the obstacle position offset exceeds 0.5 meters or the volume changes more than 20%, the area and its surrounding 1-meter range are marked as a dynamic obstacle avoidance area, and the region boundary coordinates and detection time are recorded. When obstacle avoidance is needed, the system first obtains all adjacent shelf channels reachable from the current position, and excludes the channels marked as temporarily occupied and the channels intersecting with the dynamic obstacle avoidance area through state checking. In the remaining available channels, the optimal obstacle avoidance path is selected according to the priority rules of being closest to the target point, having the largest channel width, and having the smallest angle with the current motion direction. After selecting the obstacle avoidance path, the system generates a motion trajectory considering a 0.5-meter safety margin. During the execution of obstacle avoidance, the system continuously monitors the environmental changes, and updates the path planning in time when a new dynamic obstacle is detected. This obstacle avoidance strategy based on multi-sensor fusion can effectively identify and respond to static occupation and dynamic obstacles in the shelf environment, and ensure the safe operation of the UAV.
[0118] In some embodiments, the method further comprises:
[0119] The angular velocity component in the inertial motion data is extracted, and when the rate of change of the angular velocity component exceeds a preset angular velocity threshold, the unmanned aerial vehicle is identified to be about to enter a corner scene. Within a preset time window before entering the corner scene, the fusion weight of visual positioning is adjusted to a second weight value from a first weight value through a linear gradient function, the fusion weight of inertial positioning is adjusted to a fourth weight value from a third weight value, and the fusion weight of ToF-based geometric positioning is adjusted to a sixth weight value from a fifth weight value, the second weight value is less than the first weight value, the fourth weight value is greater than the third weight value, and the sixth weight value is greater than the fifth weight value. The scene type identifier and the corridor width change sequence of the corresponding position of the grid unit are recorded in the pre-stored map, the scene type identifier includes a repeated texture scene identifier and a corner scene identifier, the scene type identifier of the grid unit within a preset distance range in front of the current position and the flight direction is obtained during flight, and the sensor fusion weight adjustment of the corresponding scene is performed in advance according to the scene type identifier in front.
[0120] The system first extracts the three-axis angular velocity [ωx, ωy, ωz] from the IMU data and calculates the rate of change dω / dt. When the rate of change of the angular velocity of any axis exceeds the preset threshold (such as 45° / s²), the system determines that the unmanned aerial vehicle is about to enter a corner scene. After the corner scene recognition trigger, the system performs dynamic weight adjustment within a 1-second time window before entering the corner: the visual positioning weight is linearly reduced from 0.5 (first weight value) to 0.2 (second weight value), because image feature matching is prone to failure at corners; the inertial positioning weight is increased from 0.3 (third weight value) to 0.5 (fourth weight value) to provide continuous motion estimation; and the ToF-based geometric positioning weight is increased from 0.2 (fifth weight value) to 0.3 (sixth weight value) to enhance the dependence on wall structure features. The weight adjustment is implemented using a linear interpolation function w(t) = w1 + (w2-w1) x t / T, where t is the current time and T is the length of the time window. At the same time, the system uses a grid storage structure to record environment information in the pre-stored map: each grid unit (such as 0.5m x 0.5m) contains a scene type identifier (1 represents a repeated texture scene and 2 represents a corner scene) and a corridor width change sequence {w1, w2,..., wn} collected at this position. n} During flight, the system calculates the grid unit index within a 5-meter range in front according to the current position [x, y] and heading angle θ, and queries the scene type identifier of these grid units. When a specific type of scene is detected in front, the corresponding weight adjustment strategy is started in advance to realize scene adaptation of the sensor fusion scheme. This scene prediction-based weight adjustment mechanism significantly improves the robustness and accuracy of multi-sensor fusion positioning in complex environments.
[0121] The indoor positioning system of the micro unmanned plane in the embodiment of the present application is described from the perspective of hardware processing below. Please refer to Figure 3 FIG. 1 is a schematic diagram of an entity device structure of the indoor positioning system of the micro unmanned plane in the embodiment of the present application.
[0122] It should be noted that Figure 3 The structure of the indoor positioning system of the micro unmanned plane shown is only an example and should not bring any limitation to the function and use range of the embodiment of the present application.
[0123] As Figure 3 shown, the indoor positioning system of the micro unmanned plane includes a central processing unit (CPU) which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) or the program loaded from the storage part 308 into the random access memory (RAM), such as performing the method described in the above embodiment. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, the ROM 302 and the RAM 303 are connected to each other through the bus 304. The input / output (I / O) interface is also connected to the bus 304.
[0124] The following components are connected to the I / O interface 305: the input part 306 including an audio input device, a button switch and the like; the output part 307 including a liquid crystal display (LCD) and an audio output device, an indicator light and the like; the storage part 308 including a hard disk and the like; and the communication part 309 including a network interface card such as a LAN (Local Area Network) card, a modem and the like. The communication part 309 performs communication processing via a network such as the Internet. The drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory and the like is installed in the drive 310 as needed so that the computer program read therefrom is installed in the storage part 308 as needed.
[0125] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program in accordance with embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising computer programs for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the present application are executed.
[0126] Note that specific examples of the computer readable storage medium can include one or more of a volatile memory, a non-volatile memory, a magnetic storage medium, an optical storage medium, or a portable storage medium. In addition, the computer readable storage medium can also be any tangible medium that is suitable for storing or carrying a program. Specifically, the computer readable storage medium of the present application can be any tangible medium that can be used to store or carry the program codes of the application and that can be accessed by a general purpose or special purpose computer system.
[0127] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved.
[0128] In particular, the indoor positioning system of the micro unmanned vehicle according to the embodiment includes a processor and a memory, and the memory stores a computer program, and the computer program is executed by the processor to implement the indoor positioning method of the micro unmanned vehicle in the ultrasonic environment.
[0129] As another aspect, the present application also provides a computer readable storage medium, which can be included in the indoor positioning system for the micro unmanned aerial vehicle as described in the above embodiments, or can exist independently without being assembled into the indoor positioning system for the micro unmanned aerial vehicle. The storage medium carries one or more computer programs, which, when executed by a processor of the indoor positioning system for the micro unmanned aerial vehicle, enable the indoor positioning system for the micro unmanned aerial vehicle to implement the indoor positioning method for the micro unmanned aerial vehicle in the ultrasonic environment as provided in the above embodiments.
[0130] The above-described embodiments are merely used to illustrate the technical solutions of the present application, but not for limiting the present application; even though the present application has been described in detail with reference to the foregoing embodiments, those ordinarily skilled in the art should understand: the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and the modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0131] In the above embodiments, according to the context, the term "when" can be interpreted as meaning "if" or "after" or "in response to determining" or "in response to detecting". Similarly, according to the context, the phrase "upon determining" or "if detecting (the stated condition or event)" can be interpreted as meaning "if determining" or "in response to determining" or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)".
[0132] Those ordinarily skilled in the art can understand that all or part of the processes in the above-described embodiments can be instructed by a computer program to relevant hardware, and the program can be stored in a computer readable storage medium, and when executed, can include the processes of the above-described embodiments. The foregoing storage medium includes ROM or random access memory (RAM), magnetic disc or optical disc, and various program code storage media.
Claims
1. An indoor positioning method of a micro unmanned aerial vehicle using ultrasonic environment perception, characterized in that, The method is applied to an indoor positioning system of a micro unmanned aerial vehicle, and comprises the following steps: Collecting image sequence data of a vision camera, inertial motion data of an inertial measurement unit (IMU), and distance measurement data of a time-of-flight (ToF) ranging sensor; Performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks, measuring left and right side wall distances by using the ToF ranging sensor, and identifying a current scene as a repetitive texture scene when a variation amplitude of the periodic spectral peaks and the left and right side wall distances is within a preset threshold range; In the repetitive texture scene, extracting a frame-to-frame displacement transformation amount between adjacent image frames in the image sequence data, and accumulating the frame-to-frame displacement transformation amount along a flight direction to obtain a visual accumulated displacement; Continuously measuring left and right side wall distances at multiple measurement positions along the flight direction by using the ToF ranging sensor, and sequentially grouping side wall distances at the multiple measurement positions to obtain a corridor width variation sequence; Performing sequence feature matching on the corridor width variation sequence and a reference width variation sequence in a pre-stored map, determining a longitudinal positioning position according to a position of the matched reference width variation sequence in the map, and calculating a lateral offset amount according to the left and right side wall distances; Inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset amount, and the inertial motion data into a filter for fusion, and outputting a fusion positioning result.
2. The method of claim 1, wherein, The step of performing frequency domain transformation on the image sequence data to obtain periodic spectral peaks, measuring left and right side wall distances by using the ToF ranging sensor, and identifying a current scene as a repetitive texture scene when a variation amplitude of the periodic spectral peaks and the left and right side wall distances is within a preset threshold range, specifically comprises the following steps: Selecting a preset number of image frames from the image sequence data to form an image analysis window, extracting feature points from each image frame in the image analysis window, and calculating a feature point matching number between adjacent image frames; When the feature point matching number is greater than a first preset threshold, performing two-dimensional fast Fourier transformation on the image frames in the image analysis window, extracting spectral components in horizontal and vertical directions in a transformed frequency domain space, and calculating peak values of the spectral components as the periodic spectral peaks; Obtaining a left and right side wall distance measurement sequence within a preset time length from the ToF ranging sensor, and calculating a mean square error of the left and right side wall distance measurement sequence as a variation amplitude of the left and right side wall distances; When the periodic spectral peaks are greater than a second preset threshold and the variation amplitude of the left and right side wall distances is less than a third preset threshold, determining that the current scene is the repetitive texture scene.
3. The method of claim 1, wherein, The step of performing sequence feature matching on the corridor width variation sequence and a reference width variation sequence in a pre-stored map, and determining a longitudinal positioning position according to a position of the matched reference width variation sequence in the map, specifically comprises the following steps: Comparing the corridor width variation sequence and the reference width variation sequence by using a sliding window, and calculating a sequence correlation coefficient at each sliding window position; acquire a sliding window position with the largest correlation coefficient, and take a starting position of the reference width variation sequence corresponding to the sliding window position as the longitudinal positioning position; calculate a distance from the current position to a corridor center line according to the left and right side wall distance, and take the distance as the lateral offset.
4. The method of claim 1, wherein, The step of inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset and the inertial motion data into a filter for fusion and outputting a fusion positioning result specifically includes: extracting an acceleration component and an angular velocity component from the inertial motion data, performing zero offset correction and integral operation on the acceleration component and the angular velocity component to obtain an inertial position prediction value; constructing a state vector, constructing a position observation equation with the inertial position prediction value, the visual accumulated displacement and the longitudinal positioning position, and the state vector including position, velocity and attitude components; constructing a lateral constraint equation with the lateral offset, and constructing an extended Kalman filter model according to the state vector, the position observation equation and the lateral constraint equation; iteratively updating the state vector by using the extended Kalman filter model to obtain the fusion positioning result.
5. The method of claim 1, wherein, After the step of inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset and the inertial motion data into a filter for fusion and outputting a fusion positioning result, the method further includes: acquiring distance measurement data of a time-of-flight ranging sensor ToF in a horizontal plane, and judging the width and spacing of a shelf channel in the repetitive texture scene; extracting a shelf edge feature from the image sequence data, and determining an extension direction of the shelf channel in combination with the distance measurement data of the time-of-flight ranging sensor ToF; when the time-of-flight ranging sensor ToF detects a front obstacle, calculating a position of an adjacent channel according to the width and spacing of the shelf channel; controlling the unmanned aerial vehicle to turn to the adjacent channel based on the visual accumulated displacement, and maintaining the relative position with the shelf during the turning process by using the shelf edge feature; recording position information of the shelf channel where the obstacle is located into the pre-stored map.
6. The method of claim 5, wherein, After the step of recording the position information of the shelf channel where the obstacle is located into the pre-stored map, the method further includes: calculating a width difference value of two adjacent shelf channels based on the distance data measured by the time-of-flight ranging sensor ToF; when the width difference value is greater than a preset width threshold, determining that a shelf with smaller width in the two adjacent shelf channels is a temporarily occupied channel; extracting a spatial distribution feature of the obstacle from the image sequence data and the distance measurement data of the time-of-flight ranging sensor ToF; when a change in the spatial distribution feature is detected, marking a change region as a dynamic obstacle avoidance region; in the process of executing an obstacle avoidance trajectory, selecting a shelf channel that is not marked as a temporarily occupied channel and a dynamic obstacle avoidance region as a candidate obstacle avoidance path.
7. The method of claim 1, wherein, After the step of inputting the visual accumulated displacement, the longitudinal positioning position, the lateral offset and the inertial motion data into a filter for fusion and outputting a fusion positioning result, the method further includes: extracting an angular velocity component in the inertial motion data, and identifying that the UAV is about to enter a corner scene when a rate of change of the angular velocity component exceeds a preset angular velocity threshold; adjusting, within a preset time window before entering the corner scene, a fusion weight of visual positioning from a first weight value to a second weight value through a linear gradient function, adjusting a fusion weight of inertial positioning from a third weight value to a fourth weight value, and adjusting a fusion weight of the ToF-based geometric positioning from a fifth weight value to a sixth weight value, the second weight value being less than the first weight value, the fourth weight value being greater than the third weight value, and the sixth weight value being greater than the fifth weight value; recording, in a grid cell of the pre-stored map, a scene type identifier and a corridor width change sequence corresponding to a position of the grid cell, the scene type identifier including a repeated texture scene identifier and a corner scene identifier; querying, according to a current position and a flight direction, a grid cell within a preset distance range in front to obtain a front scene type identifier, and performing sensor fusion weight adjustment for a corresponding scene in advance according to the front scene type identifier.
8. A micro unmanned aerial vehicle indoor positioning system, characterized in that, The indoor positioning system for the micro UAV includes one or more processors and a memory; the memory is coupled to the one or more processors, the memory is configured to store computer program code including computer instructions, and the one or more processors are configured to invoke the computer instructions to cause the indoor positioning system for the micro UAV to perform the method according to any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, The instructions, when executed on the indoor positioning system for the micro UAV, cause the indoor positioning system for the micro UAV to perform the method according to any one of claims 1-7.
10. A computer program product, characterised in that, The computer program product, when executed on the indoor positioning system for the micro UAV, causes the indoor positioning system for the micro UAV to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Unmanned aerial vehicle vision-inertia fusion indoor positioning method
CN111024066A
AGV natural scene positioning method and system based on 3D laser radar
CN120085317A