A method and device for UAV pose localization based on visual inertial odometry.

By collecting inertial data and image frame sets at the same timestamp, visual inertial odometry is used for UAV pose positioning, which solves the problems of large cumulative pose positioning errors and poor robustness of UAVs under varying lighting and occlusion environments, and achieves high-precision and robust pose positioning.

CN121453036BActive Publication Date: 2026-04-07BEIJING JIRUIXIANG AVIATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing UAV pose positioning technologies, monocular or binocular cameras are sensitive to changes in lighting. In occluded environments, the quality of feature extraction decreases, resulting in large error accumulation, poor robustness, and difficulty in accurately synchronizing visual and inertial features in time and space, as well as the inability to adaptively allocate weights.

Method used

Inertial datasets and scene image frames at the same timestamp are collected to generate a target image frame sequence. Inertial feature alignment and multimodal weighted fusion are performed to generate a fused feature sequence. Pose transformation is then performed to generate the pose localization result of the UAV camera.

Benefits of technology

It reduces the error of drone camera pose positioning, improves the robustness of pose positioning, ensures spatiotemporal synchronization through data source, achieves adaptive feature-level fusion, reduces error accumulation, and improves positioning accuracy in environments with changing lighting and occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121453036B_ABST
    Figure CN121453036B_ABST
Patent Text Reader

Abstract

This disclosure presents a method and apparatus for UAV pose localization based on visual inertial odometry. One specific implementation of the method includes: acquiring a preset inertial dataset and a set of scene image frames captured by a UAV camera at the same timestamp; generating a target scene image frame sequence based on the scene image frame set; aligning the target scene image frame sequence with inertial features according to the preset inertial dataset to generate an aligned visual feature set and an inertial motion vector set; performing multimodal weighted fusion on the aligned visual feature set and the inertial motion vector set to generate a fused feature sequence; performing pose transformation on the fused feature sequence to generate a set of pose transformations for adjacent frames; and generating a pose localization result for the UAV camera based on the set of pose transformations for adjacent frames. This implementation reduces the error in UAV camera pose localization and improves the robustness of pose localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and specifically to a method and apparatus for UAV pose localization based on visual inertial odometry. Background Technology

[0002] With the rapid development of intelligent mobile robots, self-driving cars, drones, and wearable devices, environmental perception and autonomous positioning technologies have become crucial foundations for intelligent navigation. Drone pose localization based on visual-inertial odometry (VIO) is a technology for locating the pose of a drone. Currently, the common method for locating the pose of a drone is to acquire image sequences using monocular or binocular cameras, extract feature points, match features, and solve geometric constraints to achieve pose localization.

[0003] However, when using the above method to determine the pose of a drone, the following technical problems often arise:

[0004] Using monocular or binocular cameras presents challenges due to their sensitivity to changes in lighting and the deterioration in feature extraction quality under occlusion, leading to significant cumulative errors in drone camera pose localization. Furthermore, the complex environment in which scene image frames are acquired makes it difficult to achieve precise spatiotemporal synchronization between visual and inertial features, and the inability to adaptively adjust the weight distribution between visual and inertial features according to scene changes results in poor robustness of pose localization. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose a method and apparatus for UAV pose localization based on visual inertial odometry to solve the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a method for UAV pose localization based on visual inertial odometry. The method includes: acquiring a preset inertial dataset and a scene image frame set captured by a UAV camera at the same timestamp, wherein the scene image frame set is a collection of image frames captured by a UAV at high altitude along a preset trajectory, and the acquisition time of the scene image frame set is synchronized with the preset inertial dataset; generating a target scene image frame sequence based on the scene image frame set, wherein the target scene image frame sequence is a sequence of scene image frames ordered chronologically; and performing pose localization based on the preset inertial dataset. The scene image frame sequence is aligned with inertial features to generate an aligned visual feature set and an inertial motion vector set, wherein the dimension of the aligned visual features in the aligned visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set; the aligned visual feature set and the inertial motion vector set are fused using multimodal weighted fusion to generate a fused feature sequence; the fused feature sequence is then transformed into a pose transformation set to generate a set of pose transformations for adjacent frames, wherein the set of pose transformations for adjacent image frames represents the set of pose transformations for adjacent image frames; and a pose localization result for the UAV camera is generated based on the set of pose transformations for adjacent frames.

[0008] Secondly, some embodiments of this disclosure provide a UAV pose positioning device based on visual inertial odometry. The device includes: a data acquisition unit configured to acquire a preset inertial dataset and a scene image frame set captured by a UAV camera at the same timestamp, wherein the scene image frame set is a collection of image frames captured by a UAV at high altitude along a preset trajectory, and the acquisition time of the scene image frame set is synchronized with the preset inertial dataset; a first generation unit configured to generate a target scene image frame sequence based on the scene image frame set, wherein the target scene image frame sequence is a sequence of scene image frames ordered chronologically; and an alignment unit configured to align the target scene image frame sequence according to the preset inertial dataset. A scene image frame sequence is inertial feature alignment is performed to generate an aligned visual feature set and an inertial motion vector set, wherein the dimension of the aligned visual features in the aligned visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set; a fusion unit is configured to perform multimodal weighted fusion on the aligned visual feature set and the inertial motion vector set to generate a fused feature sequence; a transformation unit is configured to perform pose transformation on the fused feature sequence to generate an adjacent frame pose transformation set, wherein the adjacent frame pose transformation set represents the set of pose transformations of adjacent image frames; a second generation unit is configured to generate a pose localization result for the UAV camera based on the adjacent frame pose transformation set.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above embodiments of this disclosure have the following beneficial effects: the UAV pose localization method based on visual inertial odometry according to some embodiments of this disclosure reduces the camera pose localization error of the UAV and improves the robustness of pose localization. Specifically, the reason for the large accumulation of camera pose localization errors and poor robustness of the UAV is that: using monocular or binocular cameras has the problem of sensitivity to changes in lighting and reduced feature extraction quality in occluded environments, resulting in a large accumulation of camera pose localization errors. Furthermore, because the environment for acquiring scene image frames is relatively complex, it is difficult to achieve accurate spatiotemporal synchronization between visual features and inertial features, and it is impossible to adaptively adjust the weight distribution between visual features and inertial features according to scene changes, resulting in poor robustness of pose localization. Based on this, the UAV pose localization method based on visual inertial odometry according to some embodiments of this disclosure first acquires a preset inertial dataset and a scene image frame set captured by the UAV camera at the same timestamp, wherein the scene image frame set is a collection of image frames captured by the UAV at high altitude along a preset trajectory, and the acquisition time of the scene image frame set is synchronized with the preset inertial dataset. Therefore, spatiotemporal synchronization can be guaranteed at the data source, avoiding the problem of misaligned data timestamps. Then, based on the aforementioned scene image frame set, a target scene image frame sequence is generated, where the target scene image frame sequence is a sequence of scene image frames ordered chronologically. This normalizes the unordered scene image frame set into a sequence, facilitating subsequent processing. Next, based on the aforementioned preset inertial dataset, the target scene image frame sequence is aligned with inertial features to generate an aligned visual feature set and an inertial motion vector set. The dimension of the aligned visual features in the aligned visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set. This achieves deep fusion and alignment of visual and inertial information at the feature level, and the identical dimension also reduces the error in the UAV's camera pose localization. Then, multimodal weighted fusion is performed on the aligned visual feature set and the inertial motion vector set to generate a fused feature sequence. This achieves adaptive and robust feature-level fusion, which can adaptively adjust the weight distribution between visual and inertial features according to scene changes, improving the robustness of pose localization. Secondly, the fused feature sequence is subjected to pose transformation to generate a set of pose transformations for adjacent frames, where the set represents the collection of pose transformations of adjacent image frames. This allows the fused high-level features to be decoded into specific geometric motion quantities. Finally, based on the set of pose transformations for adjacent frames, the pose localization result for the UAV camera is generated. This integrates discrete inter-frame motion estimations into a continuous and smooth global pose. Therefore, it effectively overcomes the problems of sensitivity to illumination changes and feature extraction quality degradation in occluded environments, reducing the accumulation of errors in UAV camera pose localization.It can adaptively adjust the weight distribution between visual and inertial features, thus improving the robustness of pose localization. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a schematic diagram of an application scenario of the UAV pose localization method based on visual inertial odometry, which is one of the embodiments of this disclosure.

[0014] Figure 2 This is a flowchart of some embodiments of the UAV pose localization method based on visual inertial odometry according to the present disclosure;

[0015] Figure 3 This is a structural schematic diagram of some embodiments of the UAV pose positioning device based on visual inertial odometry according to the present disclosure;

[0016] Figure 4 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure;

[0017] Figure 5 This is a sequence of target scene image frames according to some embodiments of the UAV pose localization method based on visual inertial odometry disclosed herein. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0023] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] Figure 1 This is a schematic diagram of an application scenario of the UAV pose localization method based on visual inertial odometry, which is one of the embodiments of this disclosure.

[0025] exist Figure 1 In the application scenario, firstly, a preset inertial dataset and a scene image frame set captured by the drone camera 101 at the same timestamp are collected. The scene image frame set is a collection of image frames captured by the drone, located at high altitude, along a preset trajectory from start point A to end point B. The acquisition time of the scene image frame set is synchronized with the preset inertial dataset; that is, at multiple time points such as T0, T1, T2 to Tn, the acquisition time of the scene image frame set is consistent with the preset inertial dataset. The dashed lines in the figure represent continuous shooting by the drone along the preset trajectory. The solid lines in the figure represent the trajectory from start point A to end point B.

[0026] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of a UAV pose localization method based on visual inertial odometry according to the present disclosure. This visual inertial odometry-based UAV pose localization method includes the following steps:

[0027] Step 201: Collect a preset inertial dataset and a set of scene image frames captured by the drone camera at the same timestamp.

[0028] In some embodiments, the execution entity (e.g., a computing device) of the UAV pose localization method based on visual inertial odometry can acquire a preset inertial dataset and a scene image frame set captured by the UAV camera at the same timestamp via a wired or wireless connection. The scene image frame set is a collection of image frames captured by the UAV at high altitude along a preset trajectory, and the acquisition time of the scene image frame set is synchronized with the preset inertial dataset.

[0029] The aforementioned timestamp can be the same UTC time under the same time system. The aforementioned drone camera can be an onboard monocular visible light camera, for example, a Sony IMX477, 4K, 30fps. The aforementioned drone camera and inertial measurement unit (IMU) are hard-connected on the same carrier board, and their clocks are uniformly triggered by a Field Programmable Gate Array (FPGA). The aforementioned preset trajectory can be a pre-set fixed trajectory, for example, a pre-set straight line trajectory between two points. The aforementioned preset trajectory can also be a pre-set non-fixed trajectory, for example, an irregular curved trajectory between two points.

[0030] The aforementioned preset inertial dataset can be a pre-defined set containing the acceleration and angular velocity of the drone's camera. For example, the preset inertial dataset could be the pre-defined "1719301204587000, -0.12, 0.05, 9.81, 0.01, -0.02, 0.03". That is, at 1719301204587ms, the drone's camera has an acceleration of -0.12m / s² on the x-axis, 0.05m / s² on the y-axis, 9.81m / s² on the z-axis, an angular velocity of 0.01rad / s on the x-axis, -0.02rad / s on the y-axis, and 0.03rad / s on the z-axis.

[0031] The aforementioned scene image frame set can be a collection of image frames containing images captured by a drone camera at the same timestamp. For example, the aforementioned scene image frame set can be a collection of image frames with a resolution of 3840×2160 captured by a drone camera at an altitude of 80m at 1719301204587ms, containing images of houses, high-voltage power line towers, vegetation, etc.

[0032] It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future wireless connection methods.

[0033] It should be noted that the aforementioned computing devices can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. For example, the computing device can be the aforementioned target terminal. When the computing device is software, it can be installed on the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.

[0034] Step 202: Based on the above scene image frame set, generate the target scene image frame sequence.

[0035] In some embodiments, the execution entity may generate a target scene image frame sequence based on the scene image frame set. The target scene image frame sequence is a sequence of scene image frames ordered chronologically.

[0036] like Figure 5 As shown, the images are sorted in chronological order as T0, T1, T2, and T3. The rectangles represent the captured scene image frames, and the dashed arrows represent the flight trajectories of the drone at different times.

[0037] As an example, the aforementioned execution entity can perform denoising processing on the aforementioned scene image frame set using wavelet transform to generate a denoised scene image frame set. Then, the denoised scene image frame set is enhanced using the Laplace operator to generate an enhanced scene image frame set. Finally, the enhanced scene image frame set is sorted according to chronological order to obtain the target scene image frame sequence.

[0038] Optionally, the aforementioned execution entity can generate a target scene image frame sequence based on the aforementioned scene image frame set through the following steps:

[0039] The first step is to synchronously crop the above scene image frame set to generate a cropped scene image frame set.

[0040] Here, the above-mentioned cropped scene image frame set can be the scene image frame set after removing the distorted or irrelevant areas of the image frames.

[0041] As an example, the aforementioned execution entity can crop the same area of ​​each scene image frame in the aforementioned scene image frame set to the same size to generate a cropped scene image frame set. For example, for a scene image frame set of size 1280×720, the edge pixels are cropped, and the 1024x640 area in the middle of the image frame is retained to remove distorted or irrelevant areas at the lens edges.

[0042] The second step is to perform resolution transformation on the above cropped scene image frame set to generate a transformed scene image frame set.

[0043] Here, the aforementioned transformed scene image frame set can be a scene image frame set after scaling the image size.

[0044] As an example, the aforementioned execution entity can scale the cropped scene image frame set to generate a transformed scene image frame set. For instance, scaling a 1024x640 image to 640x320 can generate a transformed scene image frame set. Resolution transformation can reduce subsequent computation.

[0045] The third step is to perform photometric adjustments on the above-mentioned transformed scene image frame set to generate an adjusted scene image frame set.

[0046] As an example, the aforementioned execution entity can adjust the contrast of the aforementioned transformed scene image frame set through histogram equalization and adjust the brightness of the aforementioned transformed scene image frame set through gamma correction to generate an adjusted scene image frame set.

[0047] The fourth step is to perform pixel standardization on the above-mentioned adjusted scene image frame set to generate a standardized scene image frame set.

[0048] As an example, the aforementioned execution entity can normalize the pixels of the aforementioned adjusted scene image frame set to generate a standardized scene image frame set. For example, the pixels of the aforementioned adjusted scene image frame set can be divided by 255 so that the range of pixels in the aforementioned adjusted scene image frame set is [0, 1].

[0049] The fifth step is to sort the standardized scene image frame set by timestamps to generate the target scene image frame sequence.

[0050] As an example, the aforementioned execution entity can timestamp the standardized scene image frame set according to chronological order to generate a target scene image frame sequence. For instance, the scene image frame acquired at 14:10:25 on December 11, 2025, precedes the scene image frame acquired at 14:10:45 on December 11, 2025.

[0051] Step 203: Based on the above-mentioned preset inertial dataset, perform inertial feature alignment on the above-mentioned target scene image frame sequence to generate an aligned visual feature set and an inertial motion vector set.

[0052] In some embodiments, the execution entity may perform inertial feature alignment on the target scene image frame sequence based on the preset inertial dataset to generate an aligned visual feature set and an inertial motion vector set.

[0053] Here, the aligned visual features in the aforementioned aligned visual feature set can be visual features containing UAV motion state information that are aligned with the preset inertial data timestamps in the preset inertial dataset. The aforementioned UAV motion state information may include, but is not limited to: the UAV's motion direction and the UAV's motion speed. The aforementioned inertial motion vectors in the aforementioned inertial motion vector set can be vectors containing motion features with the same dimension as the aligned visual features in the aligned visual feature set. The aforementioned motion features may include, but are not limited to: motion acceleration.

[0054] As an example, the aforementioned execution entity can use an attention mechanism to align the preset inertial dataset with the target scene image frame sequence to generate an aligned visual feature set. For instance, assuming at timestamp t=1719301204587ms: the target scene image frame is [0.12, -0.45, 0.87…0.23], the preset inertial data is [-0.12, 0.05, 9.81, 0.01, -0.02, 0.03], and the vector form of the aligned visual features might be [0.10, -0.48, 0.92…0.19]. Then, a one-dimensional convolution is performed on the target scene image frame sequence to generate a target scene image frame feature vector sequence. Finally, the target scene image frame feature vector sequence is mapped to the same dimensional space as the aligned visual feature set to obtain the inertial motion vector set.

[0055] Optionally, the aforementioned execution entity can perform inertial feature alignment on the aforementioned target scene image frame sequence based on the aforementioned preset inertial dataset through the following steps to generate an aligned visual feature set and an inertial motion vector set:

[0056] The first step is to extract visual features from the above target scene image frame sequence to generate a visual feature vector set.

[0057] Here, the visual feature vectors in the aforementioned set of visual feature vectors can be mathematical vectors that characterize the content of an image.

[0058] As an example, the aforementioned execution entity can extract visual features from the target scene image frame sequence using the Scale-Invariant Feature Transform (SIFT) algorithm to generate a visual feature vector set. Here, the aforementioned visual features can be features in the target scene image frames that can characterize the content of the scene. The scene content can include, but is not limited to, at least one of the following: urban buildings, farmland, and roads.

[0059] The second step involves performing the following alignment steps for each visual feature vector in the aforementioned visual feature vector set:

[0060] The first sub-step involves selecting the preset inertial data that is closest to the timestamp of the aforementioned visual feature vector from the preset inertial dataset.

[0061] As an example, the aforementioned execution entity can query the preset inertial data that is closest to the timestamp of the aforementioned visual feature vector from the aforementioned preset inertial dataset. For example, if the timestamp of the visual feature vector is 100.0ms, and the preset inertial dataset contains inertial data at 99.7ms and 100.2ms, the preset inertial data at 100.2ms is selected.

[0062] The second sub-step involves integrating the angular velocity of the preset duration included in the preset inertial data to generate an inertial motion vector corresponding to the aforementioned visual feature vector.

[0063] Here, the aforementioned preset duration can be the first 10ms and the last 10ms of the inertial data. The aforementioned inertial motion vector can characterize the angular change of the visual feature vector within the preset duration. The aforementioned inertial motion vector can be a three-dimensional vector, representing the rotation angle around the x-axis, y-axis, and z-axis. The aforementioned integral is a mathematical term.

[0064] The third sub-step involves concatenating and convolving the aforementioned inertial motion vectors to generate aligned visual features.

[0065] As an example, the aforementioned execution entity can concatenate the inertial motion vectors to generate a concatenated motion vector. Then, a 1×1 convolution is applied to this concatenated motion vector to generate alignment visual features. For instance, concatenating the 3D angular velocity integral vector of the inertial motion vector with the 3D mean acceleration vector of the same period yields a 6-dimensional vector, i.e., the concatenated motion vector. The dimensions of the alignment visual features are the same as the dimensions of the inertial motion vector.

[0066] Optionally, the aforementioned execution entity can perform inertial feature alignment on the aforementioned target scene image frame sequence based on the aforementioned preset inertial dataset through the following steps to generate an aligned visual feature set and an inertial motion vector set:

[0067] The first step is to perform time offset calibration on the above-mentioned preset inertial dataset to generate a zero-offset inertial dataset.

[0068] Here, the aforementioned time offset can be a tiny delay (typically a few milliseconds to tens of milliseconds) between the drone's camera exposure time and the inertial data recording time.

[0069] As an example, the aforementioned execution entity can use the Maximum CrossCorrelation (MCC) method to perform time offset calibration on the preset inertial dataset to generate a zero-offset inertial dataset. For instance, when a drone makes a rapid turn, the camera captures a scene rotation, but the inertial data is recorded 5ms too early. Through maximum crosscorrelation, it is found that when the inertial data is shifted backward by 5ms, its angular velocity change best matches the angular velocity calculated by vision. This 5ms offset is automatically compensated for, generating a time-synchronized zero-offset inertial dataset.

[0070] The second step is to perform dynamic extrinsic parameter correction on the above zero-offset inertial dataset to generate a corrected inertial dataset.

[0071] Here, the aforementioned dynamic extrinsic parameters can characterize the relative pose relationship between the UAV's camera and the zero-offset inertial dataset. The relative pose relationship can include, but is not limited to, rotation direction and angle, translation direction, and translation distance.

[0072] As an example, the aforementioned execution entity can use a sliding window optimizer (e.g., Georgia TechSmoothing and Mapping, GTSAM) to determine the optimal extrinsic parameters for the zero-offset inertial dataset. Then, the zero-offset inertial dataset is transformed according to these optimal extrinsic parameters to generate a corrected inertial dataset. For example, the optimal extrinsic parameters are multiplied by the zero-offset inertial data in the zero-offset inertial dataset to obtain the corrected inertial dataset.

[0073] The third step is to augment the corrected inertial dataset to generate an augmented inertial dataset.

[0074] As an example, the aforementioned execution entity can add zero-bias noise to the corrected inertial dataset to generate an augmented inertial dataset. This zero-bias noise changes slowly over time. This zero-bias noise can be used to simulate temperature drift. For example, by adding a zero-bias drift of 0.5° / s along the x-axis to the corrected inertial dataset, a 10g impact pulse lasting 0.1 seconds was simulated on the corrected inertial dataset.

[0075] The fourth step is to perform adaptive low-pass filtering on the augmented inertial dataset to generate a smoothed inertial dataset.

[0076] Here, the smoothed inertial dataset mentioned above can be the inertial dataset after removing high-frequency noise.

[0077] As an example, the aforementioned execution entity can use wavelet transform to perform energy distribution analysis of the signal at different scales on the aforementioned augmented inertial dataset. In response to the detection of a high-frequency useful signal (e.g., rapid turning), it automatically increases the cutoff frequency (e.g., 25Hz), and in response to the detection of a low-frequency signal (e.g., drone hovering), it automatically sets a low cutoff frequency (e.g., 5Hz) to obtain a smoothed inertial dataset.

[0078] The fifth step is to perform motion space pre-integration on the above smooth inertial dataset to generate a pre-integrated motion vector set.

[0079] Here, the pre-integrated motion vectors in the aforementioned pre-integrated motion vector set can be a compact representation of the relative motion increment over a period of time.

[0080] As an example, the aforementioned execution entity can pre-integrate the smoothed inertial dataset in the SO(3)×R³ Lie group space to generate a pre-integrated motion vector set. For example, at a time interval of t=1.0s to t=1.5s (0.5 seconds), the smoothed inertial data (100 data points) at 200Hz is integrated to obtain the pre-integrated motion vector. The pre-integrated motion vector can include: rotation increment, velocity increment, and position increment. The rotation increment can be the integrated angular velocity on the SO(3) manifold, the velocity increment can be the integrated velocity in the R³ space, and the position increment can be the double-integrated acceleration in the R³ space.

[0081] The sixth step is to interpolate the above target scene image frame sequence to generate a high temporal resolution scene image frame sequence.

[0082] As an example, the aforementioned execution entity can use motion-aware depth interpolation to interpolate the target scene image frame sequence to generate a high temporal resolution scene image frame sequence. For instance, when the angular velocity modulus of the target scene image frame exceeds a preset threshold, an interpolation mode is triggered, adding one frame every 50° / s, uniformly interpolating the target scene image frame sequence to generate a high temporal resolution scene image frame sequence. The inserted new frames are time-aligned with inertial data, and each new frame has corresponding inertial data.

[0083] The seventh step is to perform hybrid encoding on the above high temporal resolution scene image frame sequence to generate a high-dimensional visual feature set.

[0084] As an example, the aforementioned execution entity can employ a pyramid hybrid encoder to hybrid encode the high temporal resolution scene image frame sequence to generate a high-dimensional visual feature set. The pyramid hybrid encoder comprises a single-layer convolutional neural network (CNN), four Transformer layers, and a pyramid fusion layer. The CNN outputs multi-scale feature maps, while the Transformer outputs enhanced global contextual features. The pyramid fusion layer upsamples high-level semantic features and fuses them with low-level features, and performs lateral connections by adjusting the number of channels through 1×1 convolutions.

[0085] Step 8: Perform cross-attention alignment between the above high-dimensional visual feature set and the above pre-integrated motion vector set to generate an aligned visual feature set.

[0086] As an example, the aforementioned execution entity can use a bidirectional cross-attention mechanism to perform cross-attention alignment between the aforementioned high-dimensional visual feature set and the aforementioned pre-integrated motion vector set to generate an aligned visual feature set.

[0087] The ninth step is to perform balanced resampling on the above smoothed inertial dataset to generate an inertial motion vector set.

[0088] As an example, the aforementioned execution entity can map smooth inertial datasets of different dimensions to the same dimensional space to obtain a mapped inertial dataset. Then, interpolation or downsampling is performed on the time axis to resample the mapped inertial dataset to obtain a set of inertial motion vectors.

[0089] The relevant content in steps one through nine above constitutes an inventive point of this disclosure, solving the following technical problem: "loss of motion information and insufficient robustness of pose localization." Factors leading to image blurring, loss of motion information, and insufficient robustness of pose localization due to inertial feature alignment often include: interference from zero-bias drift and impact noise in real-world environments, causing image blurring and loss of motion information during rapid motion. Furthermore, in dynamic conditions such as lighting changes, motion speed, and scene complexity, in variable high-altitude environments, inertial feature alignment cannot be performed, resulting in insufficient robustness of pose localization. Solving these factors can reduce motion information loss and improve the robustness of pose localization during rapid motion. To achieve this, firstly, the aforementioned preset inertial dataset is time-offset calibrated to generate a zero-bias inertial dataset. Then, dynamic extrinsic parameter correction is performed on the zero-bias inertial dataset to generate a corrected inertial dataset. Therefore, even in real-world environments, inertial data is subject to interference such as zero-bias drift and impact noise. In the case of rapid motion, inertial feature alignment can prevent image blurring and reduce the loss of motion information. The corrected inertial dataset is augmented to generate an augmented inertial dataset. Adaptive low-pass filtering is applied to the augmented inertial dataset to generate a smoothed inertial dataset. Motion space pre-integration is performed on the smoothed inertial dataset to generate a pre-integrated motion vector set. Frame interpolation is performed on the target scene image frame sequence to generate a high temporal resolution scene image frame sequence. This improves temporal resolution through frame interpolation. Hybrid encoding is performed on the high temporal resolution scene image frame sequence to generate a high-dimensional visual feature set. This provides a high-dimensional visual feature set for subsequent processing. Cross-attention alignment is performed between the high-dimensional visual feature set and the pre-integrated motion vector set to generate an aligned visual feature set. Balanced resampling is performed on the smoothed inertial dataset to generate an inertial motion vector set. Therefore, in the case of rapid movement, inertial feature alignment reduces the loss of motion information. In the case of variable high-altitude environment, inertial feature alignment can be performed according to dynamic conditions such as changes in lighting, movement speed, and scene complexity, thereby improving the robustness of pose localization.

[0090] Step 204: Perform multimodal weighted fusion on the above-mentioned aligned visual feature set and the above-mentioned inertial motion vector set to generate a fused feature sequence.

[0091] In some embodiments, the execution entity may perform multimodal weighted fusion of the alignment visual feature set and the inertial motion vector set to generate a fused feature sequence.

[0092] As an example, the aforementioned execution entity can perform multimodal weighted fusion of the aforementioned aligned visual feature set and the aforementioned inertial motion vector set through a learnable gating network to generate a fused feature sequence.

[0093] Optionally, the aforementioned execution entity can perform multimodal weighted fusion of the aforementioned aligned visual feature set and the aforementioned inertial motion vector set through the following steps to generate a fused feature sequence:

[0094] The first step is to concatenate the above-mentioned aligned visual feature set and the above-mentioned inertial motion vector set to obtain the visual inertial feature vector set.

[0095] As an example, the aforementioned execution entity can concatenate the aligned visual feature set and the inertial motion vector set end-to-end to obtain a visual-inertial feature vector set. For instance, the aligned visual feature vector is 256-dimensional, the inertial feature vector is 256-dimensional, and the concatenation yields a 512-dimensional fused feature vector.

[0096] The second step involves inputting the aforementioned visual inertial feature vector set into a preset lightweight network to obtain the visual calibration coefficients corresponding to the aforementioned aligned visual feature set and the inertial calibration coefficients corresponding to the aforementioned inertial motion vector set.

[0097] Here, the aforementioned preset lightweight network can be a pre-defined neural network containing three fully connected layers. This preset lightweight network can also be a pre-trained network used to output calibration coefficients. Visual calibration coefficients can be weight values ​​(typically numbers between 0 and 1) used to weigh the reliability of visual information. Inertial calibration coefficients can be weight values ​​(typically numbers between 0 and 1) used to weigh the reliability of inertial information. The sum of the visual and inertial calibration coefficients is 1. For example, a visual calibration coefficient of 0.7 and an inertial calibration coefficient of 0.3 indicate that the current frame trusts the visual information more.

[0098] The third step involves weighted fusion of the alignment visual feature set and the inertial motion vector set based on the aforementioned visual calibration coefficients and inertial calibration coefficients to generate a fused feature sequence.

[0099] As an example, the aforementioned execution entity can multiply the alignment visual feature set with the visual calibration coefficients to obtain the alignment visual fusion feature sequence, and multiply the inertial motion vector set with the inertial calibration coefficients to obtain the inertial fusion feature sequence. Finally, the aforementioned alignment visual fusion feature sequence and the aforementioned inertial fusion feature sequence are added together to obtain the fusion feature sequence.

[0100] In the process of adopting technical solutions to address the problems mentioned in the background section, the following issues often arise:

[0101] Traditional channel stitching only achieves a simple combination at the data level, lacking in-depth interaction and semantic association mining between features. It cannot establish an essential connection between visual environment information and inertial motion information, and direct fusion may lead to poor robustness of pose localization.

[0102] In response to the aforementioned technical problems, the following solution was adopted:

[0103] Optionally, the aforementioned execution entity can perform multimodal weighted fusion of the aforementioned aligned visual feature set and the aforementioned inertial motion vector set through the following steps to generate a fused feature sequence:

[0104] The first step is to dynamically adjust the above inertial motion vector set to generate an adjusted inertial motion vector set.

[0105] As an example, the aforementioned execution entity can dynamically adjust the range of the aforementioned inertial motion vector set using a preset scaling factor to generate an adjusted inertial motion vector set. For instance, assuming the inertial motion vector set is [0.1, 0.2, 0.3], the preset scaling factor is 10, and the adjusted inertial motion vector set is [1.0, 2.0, 3.0], consistent with the numerical range of the visual features.

[0106] The second step is to perform channel splicing on the above-mentioned adjusted inertial motion vector set and the above-mentioned aligned visual feature set to obtain a spliced ​​feature vector set.

[0107] As an example, the aforementioned execution entity can concatenate the aforementioned adjusted inertial motion vector set and the aforementioned aligned visual feature set according to the channel to obtain a concatenated feature vector set. For example, the aligned visual feature set is [1.0, 2.0, 3.0], the adjusted inertial motion vector set is [4.0, 5.0, 6.0], and the concatenated feature vector set is [1.0, 2.0, 3.0, 4.0, 5.0, 6.0].

[0108] The third step is to perform weighted processing on the above-mentioned concatenated feature vector set to generate a weighted feature vector set.

[0109] As an example, the aforementioned execution entity can use a self-attention mechanism to perform weighted processing on the concatenated feature vector set to generate a weighted feature vector set. For example, if the concatenated feature vector set is [1.0, 2.0, 3.0, 4.0, 5.0, 6.0], the self-attention mechanism calculates the weights as [0.1, 0.2, 0.3, 0.4, 0.5, 0.6], and the weighted feature vector set is [0.1, 0.4, 0.9, 1.6, 2.5, 3.6].

[0110] The fourth step involves performing multi-scale feature extraction on the aforementioned weighted feature vector set to generate a multi-scale feature vector set. The convolution kernels between the various multi-scale feature vectors in this multi-scale feature vector set are different.

[0111] As an example, the aforementioned execution entity can perform multi-scale feature extraction on the aforementioned weighted feature vector set using convolution kernels of different scales to generate a multi-scale feature vector set. For example, one-dimensional convolution kernels of 1×1, 1×3, and 1×5 can be used. For example, if the weighted feature vector set is [0.1, 0.4, 0.9, 1.6, 2.5, 3.6], feature vectors of three scales can be extracted through multi-scale convolution: [0.1, 0.4, 0.9, 1.6, 2.5, 3.6], [0.47, 0.47, 0.97, 1.67, 2.37, 2.03], and [1.42, 1.42, 1.42, 1.42, 1.42, 1.42].

[0112] The fifth step is to perform cross-modal information interaction on the above multi-scale feature vector set to generate an interactive feature vector set.

[0113] As an example, the aforementioned execution entity can utilize the cross-attention mechanism in the Transformer architecture. This allows for cross-modal information interaction on the multi-scale feature vector set to generate an interactive feature vector set. For instance, through the cross-attention mechanism, visual features learn motion information from inertial features, and inertial features learn texture information from visual features, generating an interactive feature vector set.

[0114] The sixth step is to perform dimensionality reduction on the above-mentioned interactive feature vector set to generate a dimensionality-reduced feature set. The dimensionality of this dimensionality-reduced feature set is the same as that of the above-mentioned aligned visual feature set.

[0115] As an example, the aforementioned execution entity can use PCA or an autoencoder to perform feature dimensionality reduction on the aforementioned interactive feature vector set to generate a dimensionality-reduced feature set.

[0116] Step 7: According to the preset time step, the above-mentioned dimensionality reduction feature set is sliced ​​by a sliding window to generate a time series feature set.

[0117] Here, the preset time step can be 1.

[0118] As an example, the aforementioned execution entity can slice the dimensionality-reduced feature set using a sliding window of preset length and a preset time step to generate a time series feature set. For instance, assuming the dimensionality-reduced feature set is [1, 2, 3, 4, 5], the first window sliced ​​using a sliding window of preset length 3 and a preset time step of 1 would be [1, 2, 3], the second window [2, 3, 4], and the third window [3, 4, 5], thus obtaining the time series feature set. The numbers 1, 2, 3, 4, and 5 represent feature vectors at different times.

[0119] The eighth step is to generate fused features from the above time series feature set to produce a fused feature sequence.

[0120] As an example, the aforementioned execution entity can concatenate the time-series features in the time-series feature set in chronological order to generate the final fused feature sequence. The feature vector at each time step contains both visual and inertial information.

[0121] The content in steps one through eight above constitutes an inventive point of this disclosure, solving the technical problem of "poor robustness of pose localization." This poor robustness of pose localization often stems from the inability to establish a fundamental connection between visual environment information and inertial motion information. Traditional channel stitching only achieves a simple combination at the data level, lacking deep interaction and semantic association mining between features, thus failing to establish a fundamental connection between visual environment information and inertial motion information. Direct fusion may lead to poor robustness of pose localization. Solving these factors can improve the robustness of pose localization. To achieve this, firstly, the aforementioned inertial motion vector set is dynamically adjusted to generate an adjusted inertial motion vector set. This ensures that the two modalities participate in fusion equally in numerical terms, avoiding information imbalance. The adjusted inertial motion vector set and the aligned visual feature set are then channel-stitched to obtain a stitched feature vector set. Finally, the stitched feature vector set is weighted to generate a weighted feature vector set. Multi-scale feature extraction is performed on the aforementioned weighted feature vector set to generate a multi-scale feature vector set, wherein the convolution kernels between the various multi-scale feature vectors in the multi-scale feature vector set are different. This allows for the capture of deep interactions and semantic associations between features, establishing the essential connection between visual environment information and inertial motion information. Cross-modal information interaction is then performed on the aforementioned multi-scale feature vector set to generate an interactive feature vector set. Dimensionality reduction is then performed on the aforementioned interactive feature vector set to generate a dimensionality-reduced feature set, wherein the dimensionality of the dimensionality-reduced feature set is the same as that of the aforementioned aligned visual feature set. A sliding window slice is applied to the aforementioned dimensionality-reduced feature set according to a preset time step to generate a time-series feature set. Adaptive balance between modalities is achieved through numerical range adjustment and attention weight allocation. Finally, fusion feature generation is performed on the aforementioned time-series feature set to generate a fusion feature sequence. Therefore, the pose estimation accuracy and robustness of the UAV in complex dynamic environments are significantly improved.

[0122] Step 205: Perform pose transformation on the above fused feature sequence to generate a pose transformation set for adjacent frames.

[0123] In some embodiments, the execution entity may perform pose transformation on the fused feature sequence to generate a set of pose transformations for adjacent frames. This set of pose transformations for adjacent frames represents a collection of pose transformations for adjacent image frames.

[0124] As an example, the aforementioned execution entity can perform preset pose transformations on the fused features in the aforementioned fused feature sequence to generate a set of pose transformations for adjacent frames. The preset poses can include relative rotation and changes in camera position between adjacent frames. For example, relative rotation can be rotation angles or axis-angle vectors around the x, y, and z axes. Changes in camera position between adjacent frames can be represented by three-dimensional translation vectors, such as displacements along the x, y, and z axes.

[0125] Optionally, the aforementioned execution entity can perform pose transformation on the above-mentioned fused feature sequence through the following steps to generate a set of pose transformations for adjacent frames:

[0126] The first step is to reduce the dimensionality of the above fused feature sequence to generate a set of reduced feature vectors.

[0127] As an example, the aforementioned execution entity can use Principal Component Analysis (PCA) to reduce the dimensionality of the fused feature sequence to generate a dimensionality-reduced feature vector set.

[0128] The second step is to divide the aforementioned dimensionality-reduced feature vector set into rotation component vector sets and translation component vector sets.

[0129] Here, the rotation component vector set represents the vector part of rotation, with the direction being the rotation axis and the magnitude being radians. The translation component vector set represents the vector part of translation, with the unit being meters.

[0130] As an example, the aforementioned execution entity can divide the first three dimensions of the aforementioned dimensionality-reduced feature vector set into a rotation component vector set and the last three dimensions into a translation component vector set.

[0131] The third step is to perform an exponential mapping on the above set of rotation component vectors to generate a set of rotation matrices.

[0132] As an example, the aforementioned execution entity can perform an exponential mapping on the aforementioned set of rotation component vectors using the Rodriguez rotation formula to generate a set of rotation matrices.

[0133] The fourth step is to concatenate the above translation component vector set and the above rotation matrix set to generate a pose transformation set.

[0134] As an example, the aforementioned execution entity can combine the aforementioned set of translation component vectors and the aforementioned set of rotation matrices to generate a pose transformation set. For example, the translation component vector t = [0.5, 0.1, 0.2]. T The pose transformation T = [R, t; 0, 0, 0, 1] is formed by concatenating the rotation matrix R with the pose transformation matrix R.

[0135] The fifth step is to sort the above pose transformation sets in time order to generate pose transformation sets for adjacent frames.

[0136] As an example, the aforementioned execution entity can sort the pose transformation set according to the chronological order to generate pose transformation sets for adjacent frames.

[0137] Step 206: Generate pose localization results for the above-mentioned UAV camera based on the above-mentioned adjacent frame pose transformation set.

[0138] In some embodiments, the execution entity may generate pose localization results for the UAV camera based on the adjacent frame pose transformation set.

[0139] Optionally, the aforementioned execution entity can generate pose localization results for the aforementioned UAV camera based on the aforementioned adjacent frame pose transformation set through the following steps:

[0140] The first step is to perform chained accumulation of the above adjacent frame pose transformation sets to generate the UAV global pose sequence.

[0141] Here, the above UAV global pose sequence can be the pose (position and attitude) of each frame in the world coordinate system with the first frame as the origin.

[0142] As an example, the aforementioned execution entity can perform chained accumulation of the above adjacent frame pose transformation sets through continuous matrix multiplication to generate the UAV global pose sequence. For instance, assuming the first frame pose is the identity matrix, the adjacent frame pose transformations in the adjacent frame pose transformation sets are multiplied consecutively to obtain the UAV global pose sequence.

[0143] The second step is to perform loop closure detection on the above-mentioned UAV global pose sequence to generate a set of loop closure constraint factors.

[0144] Here, the closure constraint factors in the aforementioned set of closure constraint factors are used to correct accumulated errors. The closure constraint factor represents the constraint relationship established between the current pose and the historical pose when the UAV is detected to have returned to a previously visited location.

[0145] As an example, the aforementioned execution entity can perform loop closure detection on the UAV global pose sequence by comparing the similarity between the features of the current frame and the features of historical frames, in order to generate a set of loop closure constraint factors. For example, if the features of the 100th frame are very similar to the features of the 20th frame, then a loop is considered to have formed, and a constraint is added: "The 100th frame should be very close to the 20th frame."

[0146] The third step is to perform global graph optimization on the above-mentioned closure constraint factor set and the above-mentioned UAV global pose sequence to generate pose localization results for the above-mentioned UAV camera.

[0147] As an example, the aforementioned execution entity can use an optimization algorithm (e.g., Ceres) to perform global graph optimization on the aforementioned set of loop constraint factors and the aforementioned UAV global pose sequence to generate pose localization results for the aforementioned UAV camera. For example, the loop constraint factors in the aforementioned set of loop constraint factors indicate that frame 100 (on return) and frame 1 (starting point) are the same location. First, a loop edge is added between node 1 and node 100, constraining them to have the same pose. At this time, if frame 100 is to be aligned with frame 1, then the frames between frame 100 and frame 1 also need to be pulled accordingly. Therefore, the optimization algorithm redetermines the UAV global pose sequence, making the entire loop closed, while maintaining the correctness of the relative motion between frames as much as possible, and finally obtaining the trajectory after global graph optimization, i.e., the pose localization result.

[0148] Further reference Figure 3 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a UAV pose positioning device based on visual inertial odometry. These device embodiments are similar to... Figure 2 Corresponding to the method embodiments shown, this visual inertial odometry-based UAV pose positioning device can be specifically applied to various electronic devices.

[0149] like Figure 3As shown, a UAV pose localization device 300 based on visual inertial odometry in some embodiments includes: an acquisition unit 301, a first generation unit 302, an alignment unit 303, a fusion unit 304, a transformation unit 305, and a second generation unit 306. The acquisition unit 301 is configured to acquire a preset inertial dataset and a set of scene image frames captured by a UAV camera at the same timestamp. The scene image frame set is a collection of image frames captured by a high-altitude UAV along a preset trajectory, and the acquisition time of the scene image frame set is synchronized with the preset inertial dataset. The first generation unit 302 is configured to generate a target scene image frame sequence based on the scene image frame set. The target scene image frame sequence is a sequence of scene image frames ordered chronologically. The alignment unit 303 is configured to perform inertial feature alignment on the target scene image frame sequence according to the preset inertial dataset to generate a target scene image frame sequence. The system comprises: a visual feature set and an inertial motion vector set, wherein the dimension of the aligned visual features in the visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set; a fusion unit 304 configured to perform multimodal weighted fusion on the visual feature set and the inertial motion vector set to generate a fused feature sequence; a transformation unit 305 configured to perform pose transformation on the fused feature sequence to generate an adjacent frame pose transformation set, wherein the adjacent frame pose transformation set represents the set of pose transformations of adjacent image frames; and a second generation unit 306 configured to generate a pose localization result for the UAV camera based on the adjacent frame pose transformation set.

[0150] It is understandable that the units and references described in the visual inertial odometry-based UAV pose positioning device 300 are related. Figure 2 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the UAV pose positioning device 300 based on visual inertial odometry and the units contained therein, and will not be repeated here.

[0151] The following is for reference. Figure 4 It shows a schematic diagram of the structure of an electronic device 400 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0152] like Figure 4As shown, the electronic device 400 may include a processing unit 401 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device 400. The processing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0153] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 4 Each box shown can represent a device or multiple devices as needed.

[0154] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined above in the methods of some embodiments of this disclosure.

[0155] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0156] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0157] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a preset inertial dataset and a set of scene image frames captured by a drone camera at the same timestamp, wherein the scene image frame set is a collection of image frames captured by a drone at high altitude along a preset trajectory, and the acquisition time of the scene image frame set is synchronized with the preset inertial dataset; generate a target scene image frame sequence based on the scene image frame set, wherein the target scene image frame sequence is a sequence of scene image frames ordered chronologically; and, based on the preset inertial dataset, perform... The target scene image frame sequence is inertial feature aligned to generate an aligned visual feature set and an inertial motion vector set, wherein the dimension of the aligned visual features in the aligned visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set; the aligned visual feature set and the inertial motion vector set are multimodal weighted fused to generate a fused feature sequence; the fused feature sequence is pose transformed to generate a set of adjacent frame pose transformations, wherein the set of adjacent frame pose transformations represents the set of pose transformations of adjacent image frames; and a pose localization result for the UAV camera is generated based on the set of adjacent frame pose transformations.

[0158] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0160] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0161] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for UAV pose localization based on visual inertial odometry, characterized in that, include: Collect a preset inertial dataset and a scene image frame set captured by a drone camera at the same timestamp. The scene image frame set is a collection of image frames captured by a drone located at high altitude along a preset trajectory. The time of acquisition of the scene image frame set is synchronized with the preset inertial dataset. The preset trajectory can be a pre-set fixed trajectory or a pre-set non-fixed trajectory. Based on the scene image frame set, a target scene image frame sequence is generated, wherein the target scene image frame sequence is a scene image frame sequence ordered in chronological order; Based on the preset inertial dataset, the target scene image frame sequence is aligned with inertial features to generate an aligned visual feature set and an inertial motion vector set. The dimension of the aligned visual features in the aligned visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set. The process of aligning the target scene image frame sequence with inertial features based on the preset inertial dataset to generate the aligned visual feature set and the inertial motion vector set includes: The preset inertial dataset is calibrated with a time offset to generate a zero-offset inertial dataset; Dynamic extrinsic parameter correction is performed on the zero-offset inertial dataset to generate a corrected inertial dataset; The corrected inertial dataset is augmented to generate an augmented inertial dataset; An adaptive low-pass filter is applied to the augmented inertial dataset to generate a smoothed inertial dataset; The smooth inertial dataset is pre-integrated in motion space to generate a pre-integrated motion vector set; The target scene image frame sequence is interpolated to generate a high temporal resolution scene image frame sequence; The high temporal resolution scene image frame sequence is hybrid encoded to generate a high-dimensional visual feature set; Cross-attention alignment is performed between the high-dimensional visual feature set and the pre-integrated motion vector set to generate an aligned visual feature set; The smooth inertial dataset is resampled in a balanced manner to generate an inertial motion vector set; Multimodal weighted fusion of the alignment visual feature set and the inertial motion vector set is performed to generate a fused feature sequence, wherein the multimodal weighted fusion of the alignment visual feature set and the inertial motion vector set to generate the fused feature sequence includes: The aligned visual feature set and the inertial motion vector set are concatenated to obtain the visual inertial feature vector set; The visual inertial feature vector set is input into a preset lightweight network to obtain the visual calibration coefficients corresponding to the aligned visual feature set and the inertial calibration coefficients corresponding to the inertial motion vector set. The alignment visual feature set and the inertial motion vector set are weighted and fused according to the visual calibration coefficient and the inertial calibration coefficient to generate a fused feature sequence. The fused feature sequence is subjected to pose transformation to generate a set of pose transformations for adjacent frames, wherein the set of pose transformations for adjacent image frames represents a set of pose transformations for adjacent image frames, and wherein the pose transformation of the fused feature sequence to generate the set of pose transformations for adjacent frames includes: The fused feature sequence is dimensionality reduced to generate a dimensionality-reduced feature vector set; The reduced feature vector set is divided to generate a rotation component vector set and a translation component vector set; The set of rotation component vectors is subjected to exponential mapping to generate a set of rotation matrices; The translation component vector set and the rotation matrix set are concatenated to generate a pose transformation set; The pose transformation set is sorted in time order to generate pose transformation sets for adjacent frames; Based on the adjacent frame pose transformation set, a pose localization result for the drone camera is generated, wherein generating the pose localization result for the drone camera based on the adjacent frame pose transformation set includes: The adjacent frame pose transformation sets are chained together to generate a global pose sequence for the UAV. Loop closure detection is performed on the global pose sequence of the UAV to generate a set of loop closure constraint factors; Global graph optimization is performed on the set of loop constraint factors and the global pose sequence of the UAV to generate pose localization results for the UAV camera.

2. The method according to claim 1, characterized in that, The step of generating a target scene image frame sequence based on the scene image frame set includes: The scene image frame set is synchronously cropped to generate a cropped scene image frame set; The cropped scene image frame set is subjected to resolution transformation to generate a transformed scene image frame set; The photometric adjustment is performed on the transformed scene image frame set to generate an adjusted scene image frame set; The adjusted scene image frame set is pixel-normalized to generate a normalized scene image frame set; The standardized scene image frame set is timestamped to generate a target scene image frame sequence.

3. The method according to claim 1, characterized in that, The step of aligning the target scene image frame sequence with inertial features based on the preset inertial dataset to generate an aligned visual feature set and an inertial motion vector set includes: Visual features are extracted from the target scene image frame sequence to generate a visual feature vector set; For each visual feature vector in the set of visual feature vectors, perform the following alignment steps: Select the preset inertial data that is closest to the timestamp of the visual feature vector from the preset inertial dataset; Integrate the angular velocity of the preset duration included in the preset inertial data to generate an inertial motion vector corresponding to the visual feature vector; The inertial motion vectors are concatenated and convolved to generate aligned visual features.

4. A UAV pose positioning device based on visual inertial odometry, characterized in that, include: The acquisition unit is configured to acquire a preset inertial dataset and a scene image frame set captured by the drone camera at the same timestamp. The scene image frame set is a collection of image frames captured by the drone at high altitude along a preset trajectory. The acquisition time of the scene image frame set is synchronized with the preset inertial dataset. The preset trajectory can be a pre-set fixed trajectory or a pre-set non-fixed trajectory. The first generation unit is configured to generate a target scene image frame sequence based on the scene image frame set, wherein the target scene image frame sequence is a scene image frame sequence ordered in chronological order. An alignment unit is configured to perform inertial feature alignment on the target scene image frame sequence according to the preset inertial dataset to generate an aligned visual feature set and an inertial motion vector set, wherein the dimension of the aligned visual features in the aligned visual feature set is the same as the dimension of the inertial motion vectors in the inertial motion vector set. The step of performing inertial feature alignment on the target scene image frame sequence according to the preset inertial dataset to generate the aligned visual feature set and the inertial motion vector set includes: The preset inertial dataset is calibrated with a time offset to generate a zero-offset inertial dataset; Dynamic extrinsic parameter correction is performed on the zero-offset inertial dataset to generate a corrected inertial dataset; The corrected inertial dataset is augmented to generate an augmented inertial dataset; An adaptive low-pass filter is applied to the augmented inertial dataset to generate a smoothed inertial dataset; The smooth inertial dataset is pre-integrated in motion space to generate a pre-integrated motion vector set; The target scene image frame sequence is interpolated to generate a high temporal resolution scene image frame sequence; The high temporal resolution scene image frame sequence is hybrid encoded to generate a high-dimensional visual feature set; Cross-attention alignment is performed between the high-dimensional visual feature set and the pre-integrated motion vector set to generate an aligned visual feature set; The smooth inertial dataset is resampled in a balanced manner to generate an inertial motion vector set; The fusion unit is configured to perform multimodal weighted fusion of the alignment visual feature set and the inertial motion vector set to generate a fused feature sequence, wherein the multimodal weighted fusion of the alignment visual feature set and the inertial motion vector set to generate the fused feature sequence includes: The aligned visual feature set and the inertial motion vector set are concatenated to obtain the visual inertial feature vector set; The visual inertial feature vector set is input into a preset lightweight network to obtain the visual calibration coefficients corresponding to the aligned visual feature set and the inertial calibration coefficients corresponding to the inertial motion vector set. The alignment visual feature set and the inertial motion vector set are weighted and fused according to the visual calibration coefficient and the inertial calibration coefficient to generate a fused feature sequence. A transformation unit is configured to perform pose transformation on the fused feature sequence to generate a set of pose transformations for adjacent frames, wherein the set of pose transformations for adjacent image frames represents a set of pose transformations for adjacent image frames, and wherein performing pose transformation on the fused feature sequence to generate the set of pose transformations for adjacent frames includes: The fused feature sequence is dimensionality reduced to generate a dimensionality-reduced feature vector set; The reduced feature vector set is divided to generate a rotation component vector set and a translation component vector set; The set of rotation component vectors is subjected to exponential mapping to generate a set of rotation matrices; The translation component vector set and the rotation matrix set are concatenated to generate a pose transformation set; The pose transformation set is sorted in time order to generate pose transformation sets for adjacent frames; The second generation unit is configured to generate a pose localization result for the UAV camera based on the adjacent frame pose transformation set, wherein generating the pose localization result for the UAV camera based on the adjacent frame pose transformation set includes: The adjacent frame pose transformation sets are chained together to generate a global pose sequence for the UAV. Loop closure detection is performed on the global pose sequence of the UAV to generate a set of loop closure constraint factors; Global graph optimization is performed on the set of loop constraint factors and the global pose sequence of the UAV to generate pose localization results for the UAV camera.

5. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 3.

6. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Monocular vision inertial navigation positioning method based on self-supervised deep learning

    CN114526728A

  • Visual inertia mileage pose positioning method based on cross attention and dynamic weight

    CN120368966A