Unmanned aerial vehicle positioning method and device based on multi-scale visual inertia, and electronic equipment
By employing a multi-scale visual-inertial positioning method, utilizing camera systems and inertial measurement units with different field of view angles, the problem of positioning drift during high-speed movement and long-distance flight of UAVs was solved, achieving stable and high-precision positioning results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-21
AI Technical Summary
Under conditions of high-speed movement and long-distance flight, it is difficult for drones to maintain stable and high-precision positioning. This is mainly due to the limited number of feature point tracking attempts, the difficulty of feature registration caused by image blurring interference, and the low accuracy caused by inaccurate observations, which result in fewer map point optimization attempts.
A multi-scale visual-inertial positioning method is adopted. By equipping a camera system with different field of view angles, multi-scale image data is collected, feature analysis and fusion are performed, and combined with the inertial measurement unit, multi-scale environmental feature information is generated. Visual-inertial odometry is then calculated to determine the positioning of the UAV.
It significantly improves the positioning accuracy and stability of UAVs under high-speed movement and long-distance flight conditions. By working together with narrow-field and wide-field cameras, it enhances the stability and matching accuracy of feature points and overcomes the positioning drift problem.
Smart Images

Figure CN121904151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) navigation and control technology, and more specifically, to a UAV positioning method, device, and electronic equipment based on multi-scale visual inertial. Background Technology
[0002] Unmanned aerial vehicles (UAVs) have been widely used in surveying, field exploration, fire rescue, and agricultural production due to their advantages such as flexible flight, low cost, and no risk of personnel injury. This necessitates that UAVs solve their own localization and subsequent path planning and navigation tasks using their own sensors, even in the absence of prior environmental information. In recent years, with the rapid development of computer vision, Simultaneous Localization and Mapping (SLAM) has become one of the important technologies for UAVs to obtain their own localization. It can use sensors to estimate its own motion and simultaneously model the surrounding environment, enabling UAVs to solve their own localization problems without relying on external positioning information. However, current SLAM algorithms applied to fixed-wing UAVs suffer from long-term localization drift, meaning that the localization error increases as the distance of the UAV's flight path gradually increases.
[0003] The long-term positioning drift problem of UAVs is mainly due to the following reasons: 1. The rapid movement of fixed-wing UAVs results in fewer tracking attempts of feature points between frames, and the image blurring caused by motion leads to difficulties in feature registration; 2. Fewer observations and inaccurate inter-frame feature registration result in fewer optimization attempts for map points calculated through stereo vision or triangulation, leading to low accuracy. In summary, UAVs face the problem of difficulty in maintaining stable and high-precision positioning under conditions of high-speed movement and long-distance flight.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This invention provides a UAV positioning method, device, and electronic device based on multi-scale visual inertial, to at least solve the technical problem in the related art that UAVs are difficult to maintain stable and high-precision positioning under high-speed movement and long-distance flight conditions.
[0006] According to one aspect of the present invention, a method for locating a UAV based on multi-scale visual inertial is provided, comprising: firstly, configuring the parameters of a first camera and a second camera mounted on the UAV according to received mission requirements before the target UAV takes off, wherein the first camera is configured with a field of view angle smaller than that of the second camera; secondly, during the flight of the target UAV, using the first camera to acquire a first field of view image and using the second camera to acquire a second field of view image, thereby obtaining a multi-scale image dataset containing the first field of view image and the second field of view image; thirdly, performing feature analysis on the first field of view image and the second field of view image to obtain feature analysis results, generating multi-scale environmental feature information based on the feature analysis results, and finally determining the location information of the target UAV based on the multi-scale environmental feature information.
[0007] Further, the step of performing feature analysis on the first and second field-of-view images to obtain feature analysis results includes: extracting a first type of feature points from the first field-of-view image and extracting a second type of feature points from the second field-of-view image; performing a comparative analysis on the first and second type of feature points; if it is determined that there is a correspondence between the first and second type of feature points, fusing the first and second type of feature points to generate a unique feature point; and removing all remaining feature points that do not have a correspondence from the multi-scale image dataset to obtain a feature analysis result containing N feature points, where N is a positive integer.
[0008] Furthermore, after fusing the first type of feature points and the second type of feature points, the method further includes: for each feature point, generating a feature description vector based on the image information corresponding to the feature point in the first field-of-view image and the second field-of-view image; and performing dimensionality reduction processing on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector.
[0009] Further, the step of performing dimensionality reduction processing on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector includes: for the M-dimensional feature description vector, calculating the correlation coefficient between each element in the feature description vector and all other elements to form a correlation coefficient matrix, where M is a specified positive integer; traversing the correlation coefficient matrix to determine all element pairs corresponding to correlation coefficients lower than a preset threshold, and collecting all element pairs to obtain a low-correlation element set; based on a specified dimensionality reduction number T, selecting element pairs from the low-correlation element set, removing any element from the element pair from the feature description vector, until the feature description vector reaches the reduced dimensionality of MT, to obtain the optimized feature description vector, where T is a specified positive integer less than M.
[0010] Furthermore, before generating multi-scale environmental feature information based on the feature analysis results, the method further includes: for each feature point, fitting a quadratic surface based on the image information of the feature point in the first field-of-view image to obtain a first surface; fitting a quadratic surface based on the image information of the feature point in the second field-of-view image to obtain a second surface; parametrically processing the first surface and the second surface in the same target coordinate system to obtain a first surface function and a second surface function; defining the difference between the first surface function and the second surface function as a correlation surface function; and solving for the feature point coordinates in the target coordinate system by performing linear function analysis on the correlation surface function.
[0011] Further, the step of generating multi-scale environmental feature information based on the feature analysis results includes: determining a first intrinsic parameter applied to the first camera and a second intrinsic parameter applied to the second camera according to the received task requirements; calculating the spatial position coordinates of the feature points in three-dimensional space based on the feature point coordinates in the target coordinate system, the first intrinsic parameter, and the second intrinsic parameter; establishing a multi-scale visual feature database based on the spatial position coordinates of all the feature points; and constructing an environmental feature point cloud map using the feature description vectors corresponding to all the feature points and the spatial position coordinates to obtain the multi-scale environmental feature information, wherein the environmental feature point cloud map is used to characterize the environmental structure where the UAV flight path is located.
[0012] Further, the step of determining the positioning information of the target UAV based on the multi-scale environmental feature information includes: acquiring motion sensing data from the motion sensing device mounted inside the UAV, and acquiring an environmental feature point cloud map from the multi-scale environmental feature information; simulating the UAV pose in the environmental feature point cloud map based on the motion sensing data according to the time scale to obtain the pose information at the current time point; and applying visual inertial odometry to fuse all feature point information and all pose information in the environmental feature point cloud map within the target time period to obtain the positioning information.
[0013] According to another aspect of the present invention, a multi-scale visual-inertial unmanned aerial vehicle (UAV) positioning device is also provided, comprising: a configuration unit, configured to configure parameters of a first camera and a second camera mounted on the UAV according to received task requirements before the target UAV takes off, wherein the first camera is configured with a field of view angle smaller than that of the second camera; an acquisition unit, configured to acquire a first field of view image using the first camera and a second field of view image using the second camera during the flight of the target UAV, thereby obtaining a multi-scale image dataset containing the first field of view image and the second field of view image; an analysis unit, configured to perform feature analysis on the first field of view image and the second field of view image, obtain feature analysis results, and generate multi-scale environmental feature information based on the feature analysis results; and a determination unit, configured to determine the positioning information of the target UAV based on the multi-scale environmental feature information.
[0014] Further, the analysis unit includes: an extraction module for extracting a first type of feature points from the first field-of-view image and extracting a second type of feature points from the second field-of-view image; an analysis module for performing comparative analysis on the first type of feature points and the second type of feature points; a first fusion module for fusing the first type of feature points and the second type of feature points to generate unique feature points when it is determined that there is a correspondence between the first type of feature points and the second type of feature points; and a removal module for removing all remaining feature points that do not have a correspondence from the multi-scale image dataset to obtain a feature analysis result containing N feature points, where N is a positive integer.
[0015] Furthermore, the analysis unit further includes: a generation module, used to generate a feature description vector for each feature point after fusing the first type of feature points and the second type of feature points, based on the image information corresponding to the feature point in the first field-of-view image and the second field-of-view image; and a dimensionality reduction processing module, used to perform dimensionality reduction processing on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector.
[0016] Further, the dimensionality reduction processing module includes: a calculation submodule, used to calculate the correlation coefficient between each element in the feature description vector and all other elements for the M-dimensional feature description vector, forming a correlation coefficient matrix, where M is a specified positive integer; a traversal submodule, used to traverse the correlation coefficient matrix, determine all element pairs corresponding to correlation coefficients lower than a preset threshold, and collect all element pairs to obtain a set of low-correlation elements; and a removal submodule, used to select element pairs from the set of low-correlation elements based on a specified dimensionality reduction number T, and remove any element from the feature description vector until the feature description vector reaches the reduced MT dimensions to obtain an optimized feature description vector, where T is a specified positive integer less than M.
[0017] Furthermore, the analysis unit further includes: a fitting module, used to, before generating multi-scale environmental feature information based on the feature analysis results, fit a quadratic surface to each feature point based on the image information of the feature point in the first field-of-view image to obtain a first surface, and fit a quadratic surface based on the image information of the feature point in the second field-of-view image to obtain a second surface; a parameterization processing module, used to perform parameterization processing on the first surface and the second surface respectively in the same target coordinate system to obtain a first surface function and a second surface function; and a solution module, used to define the difference between the first surface function and the second surface function as a correlation surface function, and to solve for the coordinates of the feature points in the target coordinate system by performing linear function analysis on the correlation surface function.
[0018] Furthermore, the analysis unit further includes: a determination module, used to determine a first intrinsic parameter applied to the first camera and a second intrinsic parameter applied to the second camera based on the received task requirements; a calculation module, used to calculate the spatial position coordinates of the feature points in three-dimensional space based on the feature point coordinates in the target coordinate system, the first intrinsic parameter, and the second intrinsic parameter; an establishment module, used to establish a multi-scale visual feature database based on the spatial position coordinates of all the feature points; and a construction module, used to construct an environmental feature point cloud map using the feature description vectors corresponding to all the feature points and the spatial position coordinates, to obtain the multi-scale environmental feature information, wherein the environmental feature point cloud map is used to characterize the environmental structure where the UAV's flight path is located.
[0019] Further, the determining unit includes: an acquisition module, used to acquire motion sensing data from the motion sensing device mounted inside the UAV, and to acquire an environmental feature point cloud map from the multi-scale environmental feature information; a simulation module, used to simulate the UAV pose based on the motion sensing data in the environmental feature point cloud map according to a time scale, to obtain the pose information at the current time point; and a second fusion module, used to apply visual inertial odometry to fuse all feature point information in the environmental feature point cloud map and all the pose information within a target time period to obtain the positioning information.
[0020] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the UAV positioning method based on multi-scale visual inertial as described above.
[0021] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the UAV localization method based on multi-scale visual inertial as described above.
[0022] According to another aspect of the present invention, a computer program product is also provided, including computer instructions, wherein when the computer instructions are executed by a processor, they implement the steps of the UAV positioning method based on multi-scale visual inertial as described above.
[0023] This invention proposes a UAV localization method based on multi-scale visual inertial. First, before the target UAV takes off, the parameters of the first and second cameras on the UAV are configured according to the received mission requirements. The first camera is configured with a field of view angle smaller than that of the second camera. Then, during the flight of the target UAV, the first camera is used to acquire a first field of view image, and the second camera is used to acquire a second field of view image, resulting in a multi-scale image dataset containing the first and second field of view images. Then, feature analysis is performed on the first and second field of view images to obtain feature analysis results. Multi-scale environmental feature information is generated based on the feature analysis results. Finally, the localization information of the target UAV is determined based on the multi-scale environmental feature information.
[0024] This invention employs a multi-scale visual-inertial fusion approach. Through refined feature extraction and dimensionality reduction optimization, high-precision sub-pixel feature alignment, and an upward fusion mechanism, combined with camera systems and inertial measurement units with different field-of-view angles onboard the UAV, it achieves stable and high-precision positioning over long-distance flights. Specifically, before UAV takeoff, the dual-camera system is finely configured according to mission requirements. A narrow-field-of-view camera and a wide-field-of-view camera work collaboratively to collect multi-scale image datasets. Subsequently, feature extraction algorithms and optimization strategies are used to detect feature points and generate descriptions for the images from both fields of view. Dimensionality reduction is then performed using statistical comparison methods, significantly improving the efficiency and accuracy of feature matching. More importantly, the innovative upward fusion mechanism effectively correlates and complements the features of the narrow-field-of-view and wide-field-of-view images, greatly increasing the stability of map point observation. This overcomes the positioning drift problem caused by insufficient feature point tracking and unstable matching, thus solving the technical problem of maintaining stable and high-precision positioning for UAVs under high-speed movement and long-distance flight conditions. Attached Figure Description
[0025] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0026] Figure 1 This is a flowchart of an optional UAV localization method based on multi-scale visual inertial according to an embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram of an optional multi-scale visual input according to an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of an optional subpixel fitting quadratic surface according to an embodiment of the present invention;
[0029] Figure 4 This is a schematic diagram of an optional sub-pixel precision point according to an embodiment of the present invention;
[0030] Figure 5 This is a schematic diagram of an optional upward fusion mechanism according to an embodiment of the present invention;
[0031] Figure 6 This is a schematic diagram of an optional field-of-view image feature matching according to an embodiment of the present invention;
[0032] Figure 7 This is a schematic diagram of an optional UAV positioning device based on multi-scale visual inertial according to an embodiment of the present invention;
[0033] Figure 8 This is a structural block diagram of an electronic device for performing a multi-scale visual-inertial unmanned aerial vehicle (UAV) localization method according to an embodiment of the present invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] To facilitate understanding of the present invention by those skilled in the art, some terms or nouns involved in the various embodiments of the present invention are explained below:
[0037] Simultaneous Localization and Mapping (SLAM) is a method for robots or drones to simultaneously build an environmental map and determine their own position in an unknown environment using sensor data. In this invention, SLAM technology is extended to multi-scale visual inertial systems to improve the robustness and accuracy of localization.
[0038] An IMU, or Inertial Measurement Unit, primarily consists of an accelerometer and a gyroscope. It measures and reports an object's acceleration, angular velocity, and, under specific conditions, the direction of the Earth's magnetic field. In this invention, the IMU provides high-precision pose information over a short period to assist in the localization of visual feature points.
[0039] Subpixel feature alignment refers to reducing matching errors to the subpixel level in image feature matching through optimization algorithms, thereby improving positioning accuracy. This invention achieves subpixel-level feature point matching accuracy by fitting a quadratic surface near the feature point and utilizing the parameterization and linear equation solving of the correlation coefficient surface.
[0040] SURF, Speeded Up Robust Features, is a computer vision algorithm for image feature detection and description. It can identify image features that are scale-invariant and rotation-invariant, detect salient feature points (interest points) in an image, and describe the local information of these points by constructing feature descriptors.
[0041] Brief, Binary Robust Independent Basic Features, is a method for image feature description. It constructs a series of binary tests by randomly selecting several pairs of pixels around a feature point, comparing the gray values of these two points, and generating short, computationally fast descriptors.
[0042] Bundle Adjustment is an iterative algorithm used in visual SLAM to optimize camera pose and 3D coordinates of feature points. In this invention, fine-grained adjustment of camera pose and feature point coordinates is achieved by optimizing the overall error function.
[0043] The following embodiments of the present invention can be applied to various systems / applications / devices requiring high-precision UAV positioning and environmental mapping, enabling long-term positioning systems for fixed-wing UAVs based on multi-scale visual-inertial SLAM. The present invention uses a first camera (narrow field-of-view camera) and a second camera (wide field-of-view camera) to acquire multi-scale image data, and then processes this image data through feature extraction algorithms, dimensionality reduction optimization, sub-pixel-level feature alignment, and upward fusion mechanisms, significantly improving the positioning accuracy and stability of UAVs under high-speed movement and long-distance flight conditions.
[0044] In practical implementation, this invention first configures the parameters of the dual-camera system before the target UAV takes off, ensuring that the first camera (narrow field of view camera) and the second camera (wide field of view camera) work together to render a multi-scale image dataset that conforms to the set parameters. Then, feature point detection and descriptor generation are performed on the images of the two fields of view respectively, and the feature descriptors are then optimized by statistical comparison methods to improve the efficiency of subsequent processing. Next, sub-pixel-level feature alignment technology is used to improve the matching accuracy of feature points in images of different scales, ensuring stable tracking and high-precision registration of feature points. Next, an upward fusion mechanism is used to perform correlation fusion of feature points. By complementing the feature points of small and large field of view images, the large field of view image information is used to supplement the tracking of feature points in the narrow field of view image when tracking fails, significantly enhancing the robustness and accuracy of positioning. Finally, the iterative optimization positioning function is provided by combining the bundle adjustment method and motion sensing sensor data to achieve fine adjustment of the UAV's pose and ensure high-precision positioning during long-distance flight.
[0045] The present invention will now be described in detail with reference to various embodiments.
[0046] Example 1
[0047] According to an embodiment of the present invention, a method for UAV positioning based on multi-scale visual inertial is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0048] This invention aims to address the long-term positioning drift problem faced by existing visual inertial SLAM technology when applied to high-speed fixed-wing UAVs. Addressing the core shortcomings of high-speed UAV flight, such as fewer tracking times of feature points between image frames, image blurring, difficulty in feature registration, and consequently insufficient map point optimization, decreased accuracy, and increased cumulative error, this invention proposes a long-term UAV positioning method based on multi-scale visual inertial SLAM. The core objective is to fuse visual information at different scales (field of view) to effectively increase the tracking time and matching stability of feature points (large field of view correlation) while maintaining high-precision environmental perception capabilities (small field of view details). This significantly reduces the cumulative rate of positioning error as flight distance increases, ultimately achieving more robust and higher-precision long-term autonomous positioning for UAVs in long-distance, large-scale missions.
[0049] The implementation subject of this invention can be a UAV positioning system based on multi-scale visual inertial navigation, or integrated into various UAVs, robot navigation systems, autonomous vehicles and other intelligent mobile devices. Combining visual SLAM technology and inertial navigation technology, as well as data processing technologies such as feature point extraction, descriptor dimensionality reduction, sub-pixel-level feature alignment, and upward fusion mechanism, it provides key technical means to improve the autonomous navigation capability of UAVs in complex environments. It is particularly suitable for scenarios such as surveying and mapping, field exploration, fire rescue, and agricultural production, significantly enhancing the reliability and efficiency of UAVs when performing diverse tasks, and promoting the application and practice of UAV technology in a wider range of fields.
[0050] Specifically, in implementing this invention, firstly, a UAV system equipped with a dual-view camera is used to generate a multi-scale image dataset by setting appropriate flight and camera parameters. Then, the Surf and Brief algorithms are used to process this image data, extracting feature points and generating descriptors, followed by dimensionality reduction through statistical comparison. Next, subpixel-level feature alignment technology is used to improve the matching accuracy of feature points at different scales. Then, an upward fusion mechanism is used to correlate and fuse feature points from narrow and wide fields of view, enhancing stable tracking of feature points. Finally, bundle adjustment technology is combined to optimize the positioning results, and inertial navigation data is used for short-term positioning reinforcement, ensuring high accuracy and stability of positioning information during long-distance flight. This solves the positioning drift and accuracy degradation problems faced by UAVs during high-speed movement and long-distance flight, significantly improving the autonomous flight capability and mission execution efficiency of UAVs.
[0051] The embodiments of the present invention will now be described in detail with reference to the specific implementation steps.
[0052] Figure 1 This is a flowchart of an optional UAV localization method based on multi-scale visual-inertial imaging according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0053] Step S101: Before the target drone takes off, the parameters of the first camera and the second camera on the drone are configured according to the received mission requirements, wherein the first camera is configured with a field of view angle smaller than that of the second camera.
[0054] Specifically, mission requirements refer to the specific environmental perception and positioning accuracy requirements needed for UAVs to perform missions. For example, when performing long-distance remote sensing and mapping missions, UAVs need to maintain high-precision autonomous positioning in environments without GPS signals, while simultaneously building detailed environmental maps.
[0055] The first and second cameras on a drone refer to two camera modules with different fields of view. The first camera typically has a narrower field of view, used to capture clearer, more detailed images, suitable for precise location and identification of environmental features; the second camera has a wider field of view, covering a larger range of environmental information, enhancing the stability of feature point tracking when the drone is moving at high speed.
[0056] In another alternative embodiment, a depth camera can be used instead of the first camera. The depth camera can directly provide distance information, obtain more accurate environmental depth information, enhance the accuracy of 3D positioning of feature points, and help to build more detailed maps in complex environments. Alternatively, a panoramic camera or fisheye lens can be used instead of the second camera. The panoramic camera can capture images of the entire environment, provide a wide field of view without blind spots, further enhance the environmental perception range and the tracking stability of feature points, and is very suitable for positioning in open or dense areas.
[0057] Parameter configuration refers to the necessary settings for the two cameras before the mission begins, including but not limited to lens focal length, resolution, frame rate, shutter speed, and gain, which can be adapted in advance to specific mission environments and requirements. In addition, it also includes mission-related parameters such as the drone's flight altitude and speed, used to optimize image acquisition and environmental perception. It should be noted that the field of view (LAR) is the most important configuration target in the parameter configuration stage of this embodiment of the invention, specifically referring to the maximum visible range in the horizontal and vertical directions that the camera can capture, expressed in degrees. The LAR of the first camera is smaller than that of the second camera (e.g., the first camera is configured with a 30-degree LAR, and the second camera with a 120-degree LAR), aiming to balance image detail and coverage, and providing a multi-scale information source for the drone's positioning.
[0058] Step S102: During the flight of the target UAV, a first field-of-view image is acquired using a first camera, and a second field-of-view image is acquired using a second camera, resulting in a multi-scale image dataset containing the first and second field-of-view images.
[0059] Specifically, the first field-of-view image is image data acquired by a narrow-angle camera (e.g., a 30-degree field-of-view angle) that focuses on detailed features, while the second field-of-view image is image data acquired by a wide-angle camera (e.g., a 120-degree field-of-view angle) that covers a wider range of environmental backgrounds. Together, they constitute a multi-scale visual input. Figure 2 This is a schematic diagram of an optional multi-scale visual input according to an embodiment of the present invention, such as... Figure 2 As shown, a multi-scale image dataset is a collection of image information acquired by cameras from different field-of-view angles, resulting in multi-scale visual input. It can take into account various scenarios of UAV flight, providing a data foundation for subsequent feature analysis, matching, and map construction, and enhancing the applicability and robustness of the algorithm.
[0060] More specifically, you can first set the drone's flight altitude, intrinsic parameters of the 30-degree and 120-degree field-of-view cameras according to the actual scenario requirements, and generate a .yaml configuration file; then set the drone's long-distance flight trajectory path and save the trajectory path; based on the set parameters and the set drone flight trajectory, render the 30-degree and 120-degree field-of-view images to generate a multi-scale image dataset, such as... Figure 2 As shown.
[0061] Step S103: Perform feature analysis on the first field-of-view image and the second field-of-view image to obtain the feature analysis results, and generate multi-scale environmental feature information based on the feature analysis results.
[0062] In this embodiment of the invention, feature analysis includes the detection and description of image feature points. For example, stable feature points and descriptors (i.e., feature description vectors in this invention) are extracted from the image using SURF and Brief algorithms. These serve as the basis for visual-inertial SLAM, providing crucial information for UAV localization and environment mapping. Multi-scale environmental feature information is a comprehensive feature description encompassing environmental details and global information, formed by integrating feature points and descriptors from different viewpoint images based on the above feature analysis results. It is key to maintaining stable localization for UAVs during long-distance flights. By fusing multi-scale information, localization drift can be reduced, and the accuracy and stability of localization can be improved.
[0063] In one optional embodiment, the step of performing feature analysis on the first field-of-view image and the second field-of-view image to obtain the feature analysis result includes: extracting a first type of feature points from the first field-of-view image and extracting a second type of feature points from the second field-of-view image; performing a comparative analysis on the first type of feature points and the second type of feature points; if it is determined that there is a correspondence between the first type of feature points and the second type of feature points, fusing the first type of feature points and the second type of feature points to generate a unique feature point; and removing all remaining feature points that do not have a correspondence from the multi-scale image dataset to obtain a feature analysis result containing N feature points, where N is a positive integer.
[0064] In this embodiment of the invention, image processing algorithms (such as SURF) can be used to detect and extract significant feature points from a first field-of-view (narrow field-of-view) image and a second field-of-view (wide field-of-view) image, respectively. The narrow field-of-view image, due to its higher resolution and detail, provides a more accurate feature description; the wide field-of-view image, with its larger coverage area, captures more environmental information, helping to maintain feature stability during rapid movement. By extracting the two types of feature points separately, both detailed and extensive environmental information can be obtained.
[0065] Feature points extracted from narrow-field and wide-field images are compared. The similarity of feature descriptors is used to determine whether they describe the same physical object; for example, vector distance is used to measure the similarity of feature descriptor vectors. When the feature descriptors in two fields are sufficiently similar (below a set threshold), they are considered to represent the same environmental features, i.e., a correspondence exists. Once the correspondence of feature points is confirmed, the information of these feature points can be merged to generate a unique feature point. Specifically, this can include a weighted average of the feature point's coordinates, descriptors, and confidence scores, making the fused feature point information richer, more accurate, and better able to reflect environmental characteristics.
[0066] For feature points that cannot be matched in another view image, they may be regarded as isolated or invalid information due to occlusion, lighting changes or other factors. These feature points are marked as unreliable and will not participate in subsequent pose estimation and map construction. In order to filter out feature points that may introduce noise and ensure that the accuracy and stability of the overall localization and mapping are not affected, these unreliable feature points also need to be removed from the multi-scale image dataset to build a more accurate environment model.
[0067] This invention addresses the positioning drift problem that UAVs may encounter when performing long-distance missions. It proposes a multi-scale visual-inertial SLAM approach, employing refined feature analysis, comparative analysis and fusion, and strategies to eliminate invalid information to achieve more accurate and stable positioning for UAVs during long-distance flights. By fusing detailed information from narrow-field-of-view images with wide-field-of-view images for broader environmental perception, the goal is to enhance the system's robustness and adaptability to dynamic environments while ensuring positioning accuracy.
[0068] An optional embodiment further includes, after fusing the first type of feature points and the second type of feature points: for each feature point, generating a feature description vector based on the image information corresponding to the feature point in the first field-of-view image and the second field-of-view image; and performing dimensionality reduction processing on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector.
[0069] After feature points are confirmed to have corresponding relationships and fused, these feature points need to be described in more detail to generate feature description vectors (feature descriptors). For example, for each fused feature point, a fixed-size neighborhood region is extracted centered on that point in both the first and second field-of-view images. This region contains texture information around the feature point to understand its specific environment. A description vector can be generated based on the extracted neighborhood region using the Brief algorithm. Specifically, the Brief algorithm randomly selects a pair of pixels near the feature point and compares their brightness values to obtain a series of binary bits. These binary bits are concatenated to form the feature description vector, also known as a descriptor.
[0070] Furthermore, the generated initial feature description vectors often have high dimensionality (for example, the description vector of the Brief algorithm is usually 128-dimensional), which increases the computational burden of feature matching and may lead to processing delays, making it unsuitable for the real-time positioning needs of UAVs. In addition, high-dimensional vectors may also contain too much redundant information, affecting the efficiency and accuracy of feature matching. The purpose of the dimensionality reduction processing in this invention is to reduce the dimensionality of the feature description vector, thereby reducing the computational complexity and storage requirements of the algorithm, while retaining as much key information as possible about the feature points.
[0071] Dimensionality reduction using pre-defined statistical comparison methods involves statistical analysis and information theory principles. For example, by calculating the covariance and standard deviation among the dimensions in a descriptor, dimensions with low information content or redundancy can be identified. Taking a brief descriptor as an example, the dimensionality reduction process can specifically include the following steps:
[0072] Feature points in the image were detected using the SURF algorithm, and a 128-dimensional descriptor was generated for each feature point using the Brief algorithm. Four sample data points were selected for each dimension, and the labeled sample data points are ( ). ), ( ), ( ), ( );
[0073] The covariance of the four samples in each dimension of the descriptor is calculated using the following formula: ,in , These are the corresponding averages, and the calculation formulas are as follows: ; ;
[0074] The standard deviation of the four samples in each dimension of the descriptor is calculated using the following formula: , ;
[0075] The correlation coefficient for each sample data is calculated using the following formula: The correlation coefficient The value of is between -1 and 1, where -1 indicates a negative correlation between the two sample data, 1 indicates a positive correlation, and 0 indicates no correlation. The correlation coefficient is... The closer the value is to 0, the worse the correlation between the samples.
[0076] Compare the correlation coefficients between each sample data and remove the sample data with the weakest correlation coefficient; repeat the above steps for the 128-dimensional descriptor, delete 32 data samples, and obtain a 96-dimensional descriptor, thus achieving the dimensionality reduction effect.
[0077] In short, the calculated correlation coefficients are analyzed, and the dimensions with the weakest correlation (i.e., closest to 0) are removed. This process is repeated until the target dimension is reached. For example, reducing from 128 dimensions to 96 dimensions, this statistical comparison method can remove dimensions that contribute little or may cause noise, while retaining dimensions with high information content and good discriminative power, thereby optimizing the effectiveness and computational efficiency of the feature description vector.
[0078] In another optional embodiment, the step of performing dimensionality reduction processing on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector includes: for an M-dimensional feature description vector, calculating the correlation coefficient between each element in the feature description vector and all other elements to form a correlation coefficient matrix, where M is a specified positive integer; traversing the correlation coefficient matrix to determine all element pairs corresponding to correlation coefficients lower than a preset threshold, and collecting all element pairs to obtain a set of low-correlation elements; based on a specified dimensionality reduction number T, selecting element pairs from the set of low-correlation elements, removing any element from the element pair from the feature description vector, until the feature description vector reaches the reduced dimensionality of MT, to obtain the optimized feature description vector, where T is a specified positive integer less than M.
[0079] In this embodiment of the invention, for a feature description vector of a specific dimension M, the dimensionality reduction step involves calculating the correlation coefficient of each element in the vector relative to all other elements. This can be achieved by calculating the ratio of covariance to standard deviation one by one, ultimately forming an M×M matrix, where each value represents the degree of correlation between two elements. Next, a preset correlation coefficient threshold is set to distinguish which element pairs are lowly correlated. The threshold is chosen based on a trade-off between retaining as much information as possible and reducing computational load. The entire correlation coefficient matrix is traversed, collecting all element pairs with correlation coefficients below the preset threshold, forming a set of low-correlation elements. This step aims to identify parts of the vector with high redundancy or low information complementarity. Based on the required dimensionality reduction, element pairs are selected from the set of low-correlation elements, and one element is removed from the feature description vector. This process is repeated until the MT dimension is reached. By successively removing the least correlated elements, the feature vector is progressively optimized, ensuring that the most critical information features are retained as much as possible while reducing dimensionality.
[0080] Furthermore, before generating multi-scale environmental feature information based on feature analysis results, the process includes: for each feature point, fitting a quadratic surface based on the image information of the feature point in the first field-of-view image to obtain a first surface; fitting a quadratic surface based on the image information of the feature point in the second field-of-view image to obtain a second surface; parametrically processing the first and second surfaces in the same target coordinate system to obtain a first surface function and a second surface function; defining the difference between the first surface function and the second surface function as a correlation surface function; and solving for the feature point coordinates in the target coordinate system by performing linear function analysis on the correlation surface function.
[0081] For each extracted feature point, a series of pixels centered on that feature point are selected in both the first field-of-view image (narrow field of view) and the second field-of-view image (wide field of view). These pixels constitute a local region. In this embodiment of the invention, a quadratic surface is fitted using the least squares method, which helps to extract the location information of the feature point more accurately from the image data. Specifically, the following steps are included:
[0082] Define a subpixel-fitted quadratic surface for feature points, such as Figure 3 As shown ( Figure 3 This is a schematic diagram of an optional subpixel fitting quadratic surface according to an embodiment of the present invention. The specific method is as follows:
[0083] Let the result of the descriptor-based matching be , Located in the image middle , Located in the image middle( ), Near feature point Image patches can be fitted to quadratic surfaces The correlation coefficient surface can be defined as the difference between two surface functions, and the formula for calculating the correlation coefficient surface is: ;
[0084] The correlation coefficient surface is parameterized, and its parameterized calculation formula is as follows: ;in, Represents pixel coordinates. These are the surface parameters obtained through fitting;
[0085] Construct linear equations: ,in , It is the coefficient matrix obtained through the fitting process, i.e., substituting... , The coefficient matrix formed It corresponds The difference in pixel grayscale values;
[0086] The coefficients are obtained through SVD decomposition for the surface equation. Differentiation, the differentiation formula is: ; ;
[0087] Setting the derivative formula to zero allows us to find the minimum point of the surface equation. This yields sub-pixel precision matching feature points. Solving... The formula is: ; . Figure 4 This is a schematic diagram of an optional sub-pixel precision point according to an embodiment of the present invention, wherein... Let p be a pixel, and p is the sub-pixel precision point obtained by solving the problem.
[0088] The goal of the above steps is to achieve stable and high-precision positioning of UAVs in complex environments and long-distance flight conditions. By comprehensively applying multi-scale visual information and inertial sensor data, it aims to overcome the limitations of existing SLAM technology when fixed-wing UAVs are moving at high speeds, and improve the overall positioning performance and environmental adaptability of the system.
[0089] Furthermore, the step of generating multi-scale environmental feature information based on feature analysis results includes: determining the first intrinsic parameter applied to the first camera and the second intrinsic parameter applied to the second camera according to the received task requirements; calculating the spatial coordinates of the feature points in three-dimensional space based on the feature point coordinates in the target coordinate system, the first intrinsic parameter, and the second intrinsic parameter; establishing a multi-scale visual feature database based on the spatial coordinates of all feature points; and constructing an environmental feature point cloud map using the feature description vectors and spatial coordinates corresponding to all feature points to obtain multi-scale environmental feature information, wherein the environmental feature point cloud map is used to characterize the environmental structure in which the UAV's flight path is located.
[0090] In this embodiment of the invention, before the UAV takes off, the intrinsic parameters of the first camera (narrow field-of-view camera) and the second camera (wide field-of-view camera) are determined based on the characteristics of the upcoming mission and environmental conditions. These parameters include focal length, sensor size, and pixel size, directly affecting image quality and the accurate positioning of feature points. Using the obtained feature point coordinates in the target coordinate system, combined with the pre-set intrinsic parameters of the first and second cameras, geometric reconstruction techniques such as triangulation are used to calculate the precise position coordinates of the feature points in three-dimensional space, converting two-dimensional image information into part of a three-dimensional environment model. The three-dimensional spatial position coordinates of all feature points are stored in a database, forming a multi-scale visual feature database. This database records the environmental distribution of feature points at different scales, providing the necessary information foundation for subsequent positioning, mapping, and path planning. Using the feature description vectors and three-dimensional spatial position coordinates of all feature points, a feature point cloud map is constructed that comprehensively characterizes the environmental structure of the UAV's flight path. This cloud map includes not only the geographic coordinates of the feature points but also their appearance information and relationship with the surrounding environment, facilitating the UAV's autonomous understanding and navigation of the environment.
[0091] Alternatively, this embodiment of the invention also proposes an upward fusion mechanism. Specifically, the upward fusion mechanism is used to correlate and fuse moderate feature points in 30-degree and 120-degree field-of-view images, and to track and locate based on the feature points matched between adjacent frames and different scales. Figure 5 This is a schematic diagram of an optional upward fusion mechanism according to an embodiment of the present invention, such as... Figure 5 As shown, the specific method is as follows: feature points created from small field of view in keyframe T are... After a certain number of frames of observation and verification, the feature points are considered stable and their corresponding map points are assigned. Feature matching and association are performed on the image projected onto the large field of view. During camera movement, if a feature point can be found in the small field of view image, it is first tracked and located using the feature points in the small field of view image. If a feature point cannot be found in the small field of view image, it is tracked and located using the feature points in the large field of view image. The upward fusion mechanism effectively increases the number of map point observations and the time span in the small field of view, allowing for a larger baseline length in triangulation during local mapping, resulting in more accurate map point measurements and more precise tracking and location. When tracking fails, inertial navigation sensor information is used to recover the UAV's pose transformation.
[0092] Step S104: Determine the positioning information of the target UAV based on multi-scale environmental feature information.
[0093] Furthermore, the steps for determining the positioning information of the target UAV based on multi-scale environmental feature information include: acquiring motion sensing data from the motion sensing device mounted on the UAV, and acquiring an environmental feature point cloud map from the multi-scale environmental feature information; simulating the UAV pose in the environmental feature point cloud map based on the motion sensing data according to the time scale to obtain the pose information at the current time point; and applying the visual inertial odometry method within the target time period to fuse all feature point information and all pose information in the environmental feature point cloud map to obtain the positioning information.
[0094] It should be noted that by using multi-scale information fusion feature trajectory tracking technology to track and locate the long-distance flight trajectory of the UAV, the global perception capability of the environment by the large field-of-view camera and the detailed perception capability of the environment by the small field-of-view camera are fully combined, effectively suppressing positioning drift and improving the long-term positioning accuracy of the UAV.
[0095] Figure 6 This is a schematic diagram of an optional field-of-view image feature matching according to an embodiment of the present invention, such as... Figure 6 As shown, feature points in two consecutive frames with 30-degree and 120-degree field of view images are matched and fused. The bundle adjustment method is used, and an iterative optimization method is employed to solve for the camera pose and 3D coordinates of the feature points that minimize the overall error. The overall error formula is: ;in It is an exponential function, when three-dimensional points In multi-view frame The value is 1 when the camera has a projection point, and 0 otherwise. In the above formula... This represents the principal coordinate transformation from the main camera to other cameras, where Then it represents the first The camera intrinsic parameters of each camera are determined by finding the camera pose that minimizes the aforementioned error function. As a result of the positioning; S53: Save the above long-term positioning results as the predicted path for the UAV's long-distance flight trajectory, compare it with the algorithm's actual trajectory, and calculate the positioning error of the multi-scale visual-inertial SLAM algorithm. Finally, the saved UAV long-distance flight path can be visualized to obtain the positioning information.
[0096] Through the above steps S101 to S104, before the target UAV takes off, the parameters of the first camera and the second camera on the UAV can be configured according to the received mission requirements. The first camera is configured with a field of view angle smaller than that of the second camera. Then, during the flight of the target UAV, the first camera is used to collect the first field of view image, and the second camera is used to collect the second field of view image, so as to obtain a multi-scale image dataset containing the first field of view image and the second field of view image. Then, feature analysis is performed on the first field of view image and the second field of view image to obtain the feature analysis results. Based on the feature analysis results, multi-scale environmental feature information is generated. Finally, the positioning information of the target UAV is determined based on the multi-scale environmental feature information.
[0097] In this embodiment of the invention, a multi-scale visual-inertial fusion approach is adopted. Through refined feature extraction and dimensionality reduction optimization, high-precision sub-pixel feature alignment, and an upward fusion mechanism, combined with camera systems and inertial measurement units with different field of view angles on the UAV, a stable and high-precision positioning effect over long-distance flight is achieved. Specifically, before the UAV takes off, the dual-camera system is finely configured according to mission requirements. The narrow-field-of-view camera and the wide-field-of-view camera work together to collect multi-scale image datasets. Then, feature extraction algorithms and optimization strategies are used to detect feature points and generate descriptions for the images of the two fields of view respectively. Dimensionality reduction is then performed through statistical comparison methods, which significantly improves the efficiency and accuracy of feature matching. More importantly, the innovative upward fusion mechanism enables effective association and complementarity between narrow-field-of-view image features and wide-field-of-view image features, greatly increasing the stable observation of map points and overcoming the positioning drift problem caused by the low number of feature point tracking attempts and unstable matching. This solves the technical problem in related technologies where UAVs are unable to maintain stable and high-precision positioning under high-speed movement and long-distance flight conditions.
[0098] The invention will now be described in conjunction with another alternative embodiment.
[0099] Example 2
[0100] The UAV positioning device based on multi-scale visual inertial provided in this embodiment includes multiple implementation units, each of which corresponds to a specific implementation step in Embodiment 1 above.
[0101] Figure 7 This is a schematic diagram of an optional UAV positioning device based on multi-scale visual-inertial imaging according to an embodiment of the present invention, such as... Figure 7 As shown, the device may include: a configuration unit 71, a data acquisition unit 72, an analysis unit 73, and a determination unit 74.
[0102] The configuration unit 71 is used to configure the parameters of the first camera and the second camera carried by the target drone according to the received mission requirements before the drone takes off. The first camera is configured with a field of view angle smaller than that of the second camera.
[0103] The acquisition unit 72 is used to acquire a first field-of-view image using a first camera and a second field-of-view image using a second camera during the flight of the target UAV, thereby obtaining a multi-scale image dataset containing the first and second field-of-view images.
[0104] The analysis unit 73 is used to perform feature analysis on the first field-of-view image and the second field-of-view image, obtain the feature analysis results, and generate multi-scale environmental feature information based on the feature analysis results.
[0105] The determination unit 74 is used to determine the positioning information of the target UAV based on multi-scale environmental feature information.
[0106] The aforementioned UAV positioning device based on multi-scale visual inertial can first configure the parameters of the first and second cameras carried by the UAV before takeoff by the configuration unit 71 according to the received task requirements. The first camera is configured with a field of view angle smaller than that of the second camera. Then, during the flight of the target UAV, the acquisition unit 72 uses the first camera to acquire a first field of view image and the second camera to acquire a second field of view image, resulting in a multi-scale image dataset containing the first and second field of view images. Then, the analysis unit 73 performs feature analysis on the first and second field of view images to obtain feature analysis results and generates multi-scale environmental feature information based on the feature analysis results. Finally, the determination unit 74 determines the positioning information of the target UAV based on the multi-scale environmental feature information.
[0107] In this embodiment of the invention, a multi-scale visual-inertial fusion approach is adopted. Through refined feature extraction and dimensionality reduction optimization, high-precision sub-pixel feature alignment, and an upward fusion mechanism, combined with camera systems and inertial measurement units with different field of view angles on the UAV, a stable and high-precision positioning effect over long-distance flight is achieved. Specifically, before the UAV takes off, the dual-camera system is finely configured according to mission requirements. The narrow-field-of-view camera and the wide-field-of-view camera work together to collect multi-scale image datasets. Then, feature extraction algorithms and optimization strategies are used to detect feature points and generate descriptions for the images of the two fields of view respectively. Dimensionality reduction is then performed through statistical comparison methods, which significantly improves the efficiency and accuracy of feature matching. More importantly, the innovative upward fusion mechanism enables effective association and complementarity between narrow-field-of-view image features and wide-field-of-view image features, greatly increasing the stable observation of map points and overcoming the positioning drift problem caused by the low number of feature point tracking attempts and unstable matching. This solves the technical problem in related technologies where UAVs are unable to maintain stable and high-precision positioning under high-speed movement and long-distance flight conditions.
[0108] Furthermore, the analysis unit includes: an extraction module for extracting a first type of feature points from a first field-of-view image and a second type of feature points from a second field-of-view image; an analysis module for performing comparative analysis on the first type of feature points and the second type of feature points; a first fusion module for fusing the first type of feature points and the second type of feature points to generate unique feature points when a correspondence is determined between them; and a removal module for removing all remaining feature points that do not have a correspondence from the multi-scale image dataset to obtain a feature analysis result containing N feature points, where N is a positive integer.
[0109] Furthermore, the analysis unit also includes: a generation module, used to generate a feature description vector for each feature point based on the image information corresponding to the feature point in the first and second field-of-view images after fusing the first and second type of feature points; and a dimensionality reduction module, used to perform dimensionality reduction processing on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector.
[0110] Furthermore, the dimensionality reduction processing module includes: a calculation submodule, used to calculate the correlation coefficient between each element in the feature description vector and all other elements for an M-dimensional feature description vector, forming a correlation coefficient matrix, where M is a specified positive integer; a traversal submodule, used to traverse the correlation coefficient matrix, determine all element pairs corresponding to correlation coefficients lower than a preset threshold, and collect all element pairs to obtain a set of low-correlation elements; and a removal submodule, used to select element pairs from the set of low-correlation elements based on a specified dimensionality reduction number T, and remove any element from the feature description vector until the feature description vector reaches the reduced MT dimensions, obtaining the optimized feature description vector, where T is a specified positive integer less than M.
[0111] Furthermore, the analysis unit also includes: a fitting module, used to fit a quadratic surface to each feature point based on the image information of the feature point in the first field-of-view image before generating multi-scale environmental feature information based on the feature analysis results, to obtain a first surface, and to fit a quadratic surface based on the image information of the feature point in the second field-of-view image, to obtain a second surface; a parameterization module, used to perform parameterization processing on the first surface and the second surface in the same target coordinate system to obtain a first surface function and a second surface function; and a solution module, used to define the difference between the first surface function and the second surface function as a correlation surface function, and to obtain the coordinates of the feature points in the target coordinate system by performing linear function analysis on the correlation surface function.
[0112] Furthermore, the analysis unit also includes: a determination module, used to determine the first intrinsic parameter applied to the first camera and the second intrinsic parameter applied to the second camera according to the received task requirements; a calculation module, used to calculate the spatial position coordinates of the feature points in three-dimensional space based on the feature point coordinates in the target coordinate system, the first intrinsic parameter, and the second intrinsic parameter; an establishment module, used to establish a multi-scale visual feature database based on the spatial position coordinates of all feature points; and a construction module, used to construct an environmental feature point cloud map through the feature description vectors and spatial position coordinates corresponding to all feature points, thereby obtaining multi-scale environmental feature information, wherein the environmental feature point cloud map is used to characterize the environmental structure in which the UAV's flight path is located.
[0113] Furthermore, the determining unit includes: an acquisition module, used to acquire motion sensing data from the motion sensing device mounted inside the UAV, and to acquire an environmental feature point cloud map from multi-scale environmental feature information; a simulation module, used to simulate the UAV pose based on the motion sensing data in the environmental feature point cloud map according to the time scale, and to obtain the pose information at the current time point; and a second fusion module, used to apply the visual inertial odometry method to fuse all feature point information and all pose information in the environmental feature point cloud map within the target time period to obtain positioning information.
[0114] The aforementioned UAV positioning device based on multi-scale visual inertial can also include a processor and a memory. The configuration unit 71, acquisition unit 72, analysis unit 73, determination unit 74, etc., are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.
[0115] The aforementioned processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured. By adjusting kernel parameters, the parameters of the first and second cameras on the target UAV are configured according to the received mission requirements before takeoff. During the target UAV's flight, the first camera acquires a first field-of-view image, and the second camera acquires a second field-of-view image. Feature analysis is performed on the first and second field-of-view images, and multi-scale environmental feature information is generated based on the feature analysis results, thereby determining the target UAV's positioning information.
[0116] The aforementioned memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0117] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program with the following method steps: before the target UAV takes off, the parameters of a first camera and a second camera mounted on the UAV are configured according to the received task requirements, wherein the first camera is configured with a field of view angle smaller than that of the second camera; during the flight of the target UAV, the first camera is used to acquire a first field of view image, and the second camera is used to acquire a second field of view image, thereby obtaining a multi-scale image dataset containing the first and second field of view images; feature analysis is performed on the first and second field of view images to obtain feature analysis results, and multi-scale environmental feature information is generated based on the feature analysis results; the positioning information of the target UAV is determined based on the multi-scale environmental feature information.
[0118] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute any one of the above-described embodiments of the UAV positioning method based on multi-scale visual inertial.
[0119] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the multi-scale visual-inertial unmanned aerial vehicle (UAV) positioning method based on any one of the above embodiments.
[0120] Figure 8 This is a structural block diagram of an electronic device for performing a multi-scale visual-inertial UAV localization method according to an embodiment of the present invention, such as... Figure 8 As shown, the electronic device may include: one or more ( Figure 8 (Only one is shown) Processor 802, memory 804, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0121] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the UAV positioning method and device based on multi-scale visual inertial perception in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned UAV positioning method based on multi-scale visual inertial perception. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0122] Those skilled in the art will understand that Figure 8 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 8 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 8 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.
[0123] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0124] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0125] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0128] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0130] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A UAV localization method based on multi-scale visual-inertial, characterized in that, include: Before the target drone takes off, the parameters of the first and second cameras on the drone are configured according to the received mission requirements. The first camera is configured with a field of view angle smaller than that of the second camera. During the flight of the target UAV, the first camera is used to acquire a first field-of-view image, and the second camera is used to acquire a second field-of-view image, resulting in a multi-scale image dataset containing the first field-of-view image and the second field-of-view image. Feature analysis is performed on the first and second field-of-view images to obtain feature analysis results, and multi-scale environmental feature information is generated based on the feature analysis results. The positioning information of the target UAV is determined based on the multi-scale environmental feature information.
2. The UAV localization method based on multi-scale visual-inertial as described in claim 1, characterized in that, The steps of performing feature analysis on the first and second field-of-view images to obtain feature analysis results include: Extract a first type of feature points from the first field-of-view image, and extract a second type of feature points from the second field-of-view image; A comparative analysis was performed on the first type of feature points and the second type of feature points; If it is determined that there is a correspondence between the first type of feature points and the second type of feature points, the first type of feature points and the second type of feature points are fused to generate a unique feature point; All remaining feature points that do not have a corresponding relationship are removed from the multi-scale image dataset to obtain a feature analysis result containing N feature points, where N is a positive integer.
3. The UAV positioning method based on multi-scale visual-inertial as described in claim 2, characterized in that, After fusing the first type of feature points and the second type of feature points, the process further includes: For each feature point, a feature description vector is generated based on the image information corresponding to the feature point in the first field-of-view image and the second field-of-view image; The feature description vector is dimensionality reduced based on a preset statistical comparison method to obtain an optimized feature description vector.
4. The UAV positioning method based on multi-scale visual-inertial as described in claim 3, characterized in that, The step of performing dimensionality reduction on the feature description vector based on a preset statistical comparison method to obtain an optimized feature description vector includes: For the M-dimensional feature description vector, calculate the correlation coefficient between each element in the feature description vector and all other elements to form a correlation coefficient matrix, where M is a specified positive integer; Traverse the correlation coefficient matrix, determine all element pairs corresponding to correlation coefficients below a preset threshold, and collect all element pairs to obtain a set of low-correlation elements; Based on the specified dimensionality reduction number T, select element pairs from the set of low-correlation elements, remove any element from the feature description vector, until the feature description vector reaches the reduced MT dimension, and obtain the optimized feature description vector, where T is a specified positive integer less than M.
5. The UAV localization method based on multi-scale visual-inertial as described in claim 3, characterized in that, Before generating multi-scale environmental feature information based on the feature analysis results, the process also includes: For each feature point, a quadratic surface is fitted based on the image information of the feature point in the first field of view image to obtain a first surface, and a quadratic surface is fitted based on the image information of the feature point in the second field of view image to obtain a second surface; Parametric processing is performed on the first surface and the second surface in the same target coordinate system to obtain the first surface function and the second surface function; The difference between the first surface function and the second surface function is defined as the correlation surface function, and the coordinates of the feature points in the target coordinate system are obtained by performing linear function analysis on the correlation surface function.
6. The UAV localization method based on multi-scale visual-inertial as described in claim 5, characterized in that, The steps for generating multi-scale environmental feature information based on the feature analysis results include: Based on the received task requirements, determine the first intrinsic parameter applied to the first camera and the second intrinsic parameter applied to the second camera; Based on the feature point coordinates in the target coordinate system, the first intrinsic parameter, and the second intrinsic parameter, calculate the spatial position coordinates of the feature point in three-dimensional space; A multi-scale visual feature database is established based on the spatial coordinates of all the feature points. An environmental feature point cloud map is constructed using the feature description vectors corresponding to all the feature points and the spatial location coordinates to obtain the multi-scale environmental feature information. The environmental feature point cloud map is used to characterize the environmental structure in which the UAV's flight path is located.
7. The UAV localization method based on multi-scale visual-inertial as described in claim 1, characterized in that, The step of determining the positioning information of the target UAV based on the multi-scale environmental feature information includes: Acquire motion sensing data from the motion sensing device inside the UAV, and obtain environmental feature point cloud map from the multi-scale environmental feature information; Based on the motion perception data, the UAV pose is simulated in the environmental feature point cloud map according to the time scale to obtain the pose information at the current time point. Within the target time period, the visual inertial odometry method is applied to fuse all feature point information and all pose information in the environmental feature point cloud map to obtain the positioning information.
8. A UAV positioning device based on multi-scale visual-inertial, characterized in that, include: The configuration unit is used to configure the parameters of the first camera and the second camera carried by the target drone according to the received mission requirements before the drone takes off, wherein the first camera is configured with a field of view angle smaller than that of the second camera. The acquisition unit is used to acquire a first field-of-view image using the first camera and a second field-of-view image using the second camera during the flight of the target UAV, thereby obtaining a multi-scale image dataset containing the first field-of-view image and the second field-of-view image. The analysis unit is used to perform feature analysis on the first field-of-view image and the second field-of-view image, obtain feature analysis results, and generate multi-scale environmental feature information based on the feature analysis results. The determining unit is used to determine the positioning information of the target UAV based on the multi-scale environmental feature information.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the UAV positioning method based on any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the UAV localization method based on multi-scale visual inertial as described in any one of claims 1 to 7.